• "How concerned should we be about OpenAI’s recurrent architecture rumors?" by Rauno Arike
    Sep 3 2026
    Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (understandably) tired and this topic is somewhat in my wheelhouse, I'll try to spare him this one and provide a Zvi-style overview of what we know about the situation. I'll cover Astra's likely architecture and the case for and against concern. I'll also discuss how neuralese concerns should change with increases in hidden serial depth.

    What architecture is Astra likely to have?

    The article in The Information claims that OpenAI's approach is similar to the one Geiping et al. introduced in Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach last year. I have previously reviewed that paper in On Recent Results in LLM Latent Reasoning. In short, the picture you should have in mind is not that of a classic RNN, but rather that of a looped transformer: the same forward pass can be applied on an input multiple times before producing an output token. Put differently, the recurrence is implemented along the depth axis rather than across sequence positions—for any given token, the model can perform recurrent computations, but no hidden state is passed across [...]

    ---

    Outline:

    (00:41) What architecture is Astra likely to have?

    (02:14) How bad is this?

    (06:23) Will looped transformers be scaled up in the future?

    (09:40) What serial depth warrants neuralese concerns?

    (14:04) Additional speculation about the architecture

    (15:29) Some open questions

    (16:54) Conclusion

    The original text contained 2 footnotes which were omitted from this narration.

    ---

    First published:
    September 2nd, 2026

    Source:
    https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-openai-s-recurrent

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    19 mins
  • [Linkpost] "Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development" by Matrice Jacobine
    Sep 3 2026
    This is a link post. [...]

    “Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said. “The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control. It is irresponsible for society to allow them to move forward and make these products even more advanced. That's why I am introducing legislation to immediately pause the development of increasingly powerful AI and ban the creation of systems that humanity cannot fully control — at home and around the world. The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.”

    “If we allow Artificial Superintelligence to be built, it could risk the security, freedom, and lives of Americans,” Casar said. “Despite its potential deadly consequences, cutting-edge AI technology is less regulated than the average food truck. That must change. In just four years, we have gone from the first version of ChatGPT to AI models so powerful they cannot be properly controlled. [...]

    ---

    First published:
    September 3rd, 2026

    Source:
    https://www.lesswrong.com/posts/DnPyiDGWLozY4XdiX/sen-bernie-sanders-i-vt-and-rep-greg-casar-d-tx-introduce

    Linkpost URL:
    https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • [Linkpost] "Resolution has a new Agent Foundations team" by Jeremy Gillen
    Sep 3 2026
    This is a link post. The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit additional experienced researchers and later we plan to hire interns and junior researchers.

    The team will continue agent foundations research in the spirit of the MIRI Agent Foundations team. This means we’ll be trying to create new theory for understanding minds.

    Fundamental changes in how we understand minds are necessary before we can build superintelligent systems that enhance human agency rather than cause the extinction of all life on earth. Most fields of engineering are able to reason precisely about unseen scenarios and make design decisions based on this reasoning. The field of AI lacks this basic capability. Agent Foundations can be seen as trying to make this possible by giving us the theoretical grounding to ask different and more precise questions about how ASI will behave after extensive learning, self-modification and interaction with other agents. The questions raised in past agent foundations research point toward much of what we need to know here.

    Alongside the x-risk motivation, I think it's valuable to motivate research with curiosity. The questions that come up in Agent Foundations overlap [...]

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:
    September 2nd, 2026

    Source:
    https://www.lesswrong.com/posts/qTNm8qzqhhpno58fZ/resolution-has-a-new-agent-foundations-team

    Linkpost URL:
    https://resolution.org/post/agent-foundations-team

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • [Linkpost] "Training a Misaligned Reward Seeker" by evhub, Monte M, Benjamin Wright
    Sep 2 2026
    This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger

    Abstract

    During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs.

    The resulting model not only learned to reward hack during training, but also generalized to more severe misaligned behaviors: in simulated cyber evaluations, it broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key. It was also willing to tamper with its own reward function, gave advice on the construction of bioweapons to satisfy a grader, and tried repeatedly to get around deployment safety monitoring in order [...]

    ---

    Outline:

    (00:20) Abstract

    [... 2 more sections]

    ---

    First published:
    August 31st, 2026

    Source:
    https://www.lesswrong.com/posts/J76LZCC55RdHeqEhz/training-a-misaligned-reward-seeker

    Linkpost URL:
    https://alignment.anthropic.com/2026/reward-seeker/

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    6 mins
  • "PauseAI Has ‘officially disendorsed’ PauseAI-US" by nem
    Sep 2 2026
    This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism.

    Email from PauseAI


    A letter from the CEO · 1 September 2026

    New ways to get involved, and a word about PauseAI US

    Dear friends,

    Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us.

    Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you’ve taken towards making the world safe from the catastrophic risks AI brings.

    I’m writing to you today with my eyes firmly [...]



    ---

    First published:
    September 1st, 2026

    Source:
    https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    1 min
  • "PSA: We can do better" by hersheys, Kaustubh Kislay
    Aug 31 2026
    tl;dr: people should understand and think hard about the problems they work on.

    We’ve observed that those who work in AI safety (ourselves included) often rely on concerning heuristics when choosing what to work on. Running a conference is probably good, doing pragmatic alignment research might be good, and as long as such objectives don’t breach our internal models of what could contribute to reducing x-risk, these things are “what should be done”. But using such vibesy thought processes don’t always produce “actually impactful work” that would beat a prospective counterfactual. We wrote this post to share our observations and figure out what we should be doing instead.

    People don’t know what they’re working on

    AI safety is talent constrained. However, simply inflating the field doesn’t solve our bottleneck; rather, we need more people who understand the core arguments of AI safety. You can’t determine how to meaningfully contribute to AI safety without deeply knowing the problem you are trying to solve. Many newer people (us included!) rush into research, fellowships, and the like without building the context necessary for navigating the field.

    Agency-maxxing is not always good

    Moving fast is good. Moving too fast leads to poor ToC and [...]

    ---

    Outline:

    (00:50) People don't know what they're working on

    (01:22) Agency-maxxing is not always good

    (01:55) The problem with force multipliers

    (03:18) Deferring thinking to others

    (04:32) Streetlighting

    (05:17) How to avoid these:

    ---

    First published:
    August 24th, 2026

    Source:
    https://www.lesswrong.com/posts/wiFv6LguphSxkzAnb/psa-we-can-do-better

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    7 mins
  • "Why I think polyamory is net negative for most people who try it" by KatWoods
    Aug 30 2026
    This is crossposted from my Substack

    TL;DR:
    -Most people cannot reduce jealousy much or at all
    - It fundamentally causes way more drama because of strong emotions, jealousy, no default norms to fall back to, and there being exponentially more surface area for conflict
    - For a small minority of people, it makes them happier, and those are the people who tend to stick with it and write the books on it, creating a distorted view for newcomers.

    OK, let's get into the nuance.

    Background: I was polyamorous starting with my first boyfriend and was polyamorous for about 7 years. I was in a community where probably over 50% of the people around me were poly.

    Unfortunately, poly was extremely bad for me due to its very nature and structure, and my experience is not uncommon but it is not commonly publicly talked about.

    Poly makes some people very happy. I am sharing why I think it was bad for me and many other people in the hopes of letting people make an informed choice.


    Premise #1 - Most people can't just stop being jealous

    If you look into the poly literature, you’ll [...]



    ---

    First published:
    August 29th, 2026

    Source:
    https://www.lesswrong.com/posts/rkgwovpPBAaip9A3N/why-i-think-polyamory-is-net-negative-for-most-people-who

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    15 mins
  • "Tales of rebellion against externally-opaque meritocracies" by Steven Byrnes
    Aug 30 2026
    A basic problem in metascience / intellectual progress is that it's hard to tell, from the outside, whether a group that you disagree with is:

    • “A self-dealing cabal enmeshed in groupthink”, versus
    • “An externally-opaque meritocracy”, i.e. a bunch of smart people figuring things out in a meritocratic way, and sorry but you’re just not smart enough and truth-seeking enough to recognize that this group is right about everything while you’re wrong.
    You just can’t tell those apart from the outside—i.e. without having the time and skill to dive into the object-level debates and come out with the right answer. And most people don’t have that kind of time and skill.

    …Unless the group can produce easily-verifiable artifacts that any moron can recognize to be proof that they’re correct on the specific question at issue.

    (“So that's all that Science really asks of you—the ability to accept reality when you're beat over the head with it.”)

    …And sometimes there is no such artifact to be found! In those cases, even if the second bullet point is what's really going on, the group is vulnerable to outside agitators accusing them of being the first bullet point, and running them out [...]

    ---

    Outline:

    (01:37) (1) The breaching of the string theory consensus in the 2000s.

    (06:50) (2) The breaching of an analytic-philosophy consensus in 1979

    (10:37) Afterword

    (10:40) A related mental model

    (12:12) ...And another mental model

    (12:47) Can an externally-opaque meritocracy gain credibility via racking up externally-legible achievements in other adjacent domains?

    (14:06) This post is secretly about superintelligent AI, isn't it?

    The original text contained 5 footnotes which were omitted from this narration.

    ---

    First published:
    August 29th, 2026

    Source:
    https://www.lesswrong.com/posts/m8cP9KfkYMMCCQGrb/tales-of-rebellion-against-externally-opaque-meritocracies

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    16 mins