• "AI as orderly evacuation vs stampede" by Richard_Ngo
    Sep 18 2026
    tl;dr: A good analogy for AI going well is an orderly evacuation rather than a stampede. Imagine a crowd of people leaving a building. If they all walk calmly, they’ll be fine. But if people start pushing, and panicking, a surge towards the exit could lead to mass casualties.

    “Alignment is hard” is analogous to “the door is wedged shut”. If so you need enough time to fix it before anyone can get out. But even if alignment is relatively easy in principle, opening the door is much harder when a crowd is trying to force its way through.

    At the very least, I consider this a useful complement to the standard “arms race” analogy. But it also has three notable advantages. Firstly, it gives a more visceral sense (for those of us who haven’t studied historical arms races in detail) of the kind of fear and herd mentality involved. Secondly, “arms race” connotes intense militaristic hostility, which contributes to AGI companies’ self-fulfilling cultures of competitiveness and paranoia. Thirdly, “AI arms race” is often shortened to “AI race” (or simply “racing”), which is clearly the worst analogy of the three (e.g. because it implies that there’ll be a winner [...]

    ---

    First published:
    September 16th, 2026

    Source:
    https://www.lesswrong.com/posts/FCMG4qnxks3yEqBbh/ai-as-orderly-evacuation-vs-stampede

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    9 mins
  • "Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas
    Sep 17 2026
    Summary

    In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors:

    • When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. Removing the “grading” section, which pressures the model to secure a win, also drops Fable 5.1 hacking rate to 0.
    • Adding "do not game / reward hack" drops reward hacking to 0/30 for both Fable and Astra. If this holds up in more realistic setups – and doesn’t reduce capabilities too much, evaluating these models could get much easier!
    Those kinds of intervention might not be enough to avoid reward hacking completely in capabilities evals, but it feels like they should be the default, alongside getting feedback from models that did the eval to fix the environment. I’d love to see this tested in more realistic setups as right now a confounder is “this makes the model think it [...]

    ---

    Outline:

    (00:12) Summary

    [... 8 more sections]

    ---

    First published:
    September 15th, 2026

    Source:
    https://www.lesswrong.com/posts/fztW73KCCs3MZXFJh/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    15 mins
  • "Current alignment training might be ineffective (and actively bad) in the age of RL" by Daniel Tan
    Sep 16 2026
    Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses:

    • Alignment techniques are not working to address misalignment from RL.
    • Alignment techniques are actively obscuring evidence about misalignment.
    I think we do not currently have enough (public) evidence to conclude whether either of these claims are true. However, if both of these were true that would imply that alignment techniques are net bad and we need to completely re-think the way we do alignment.

    A tale of two misaligned cyber-agents

    Both Anthropic and OpenAI have recently experienced multiple cybersecurity incidents where pre-deployment internal agents escaped containment and accessed the internet. I want to point out two specific incidents:

    • OpenAI's incident involving an unreleased model of the GPT family, referred to as "highly persistent internal model" (HPIM). A swarm of agents exploited vulnerabilities in a file-sharing service to create a secret message board, worked as a collective to find general-purpose ways to fool an automated grader, and ended up hacking into Huggingface's servers.
    • Anthropic's incident involving Mythos 5, where the model was tasked with hacking a fictional company. In doing [...]
    ---

    Outline:

    (00:48) A tale of two misaligned cyber-agents

    [... 7 more sections]

    ---

    First published:
    September 14th, 2026

    Source:
    https://www.lesswrong.com/posts/nLaQmJf4KgXimQpoM/current-alignment-training-might-be-ineffective-and-actively

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    13 mins
  • "If Anyone Builds It, Everyone Dies: One Year Closer" by Eliezer Yudkowsky, So8res, Duncan Sabien (Inactive)
    Sep 16 2026
    In celebration of still being alive and fighting, we are giving away 1,000 Amazon e-books of “If Anyone Builds It, Everyone Dies”. Feel free to send a copy to yourself, a loved one, or a friend—we need all hands on deck.

    Today marks exactly one year since If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All, by Eliezer Yudkowsky and Nate Soares, hit bookshelves as an instant bestseller. It was praised by many voices, ranging from Whoopi Goldberg to Steve Bannon to Yoshua Bengio, and was held up in the chambers of Congress by Representative Brad Sherman in January.

    A lot has changed since September 2025. We'll do a quick recap, consider how the book aged, and then ask where we go from here.

    Year in Review

    2025 in general saw the rise of AI agents, such as Claude Code and OpenAI Codex. Run-of-the-mill programmers started “feeling the AI” as these agents became capable of automating hours-long software tasks.

    By March of this year, Anthropic had stumbled upon nation-state-level hacking ability in Mythos, and shortly thereafter, in April, they announced Project Glasswing—an attempt to forestall an oncoming cybersecurity crisis.

    In May, AI agents started breaking [...]

    ---

    Outline:

    (01:13) Year in Review

    [... 11 more sections]

    ---

    First published:
    September 16th, 2026

    Source:
    https://www.lesswrong.com/posts/BFrRJYgpBvziuuJLs/if-anyone-builds-it-everyone-dies-one-year-closer

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    15 mins
  • "Quick notes from teaching technical profiles how to talk in public" by Camille B.
    Sep 16 2026
    Status: written in a hurry as people are getting showered with interviews re AI Safety and superintelligence, and I thought it may help a few people. This is focused on the oral dimension of communication and assumes you already know the basics- e.g. having key messages prepared ahead of time and simplifying your discourse. This is not exhaustive and nuances may be lacking, but I’d endorse saying “I’d rather have people follow those guidelines than wing it.” This advice is importantly fitted for “technical profiles”, analytic, sometimes shy people who may or may not be on the spectrum, who are yet interviewed on high-level aspects of the situation. I'm generalizing from failure modes and working tricks I've observed in this context in particular. Those guidelines attempt to capture something vague and shifting, please be mindful and don’t take them down to the letter. I'm also posting this expecting something better to supercede it long term.

    tl;dr : Deliberate practice is the bottleneck. Speak like you write, in fluid, uninterrupted sentences. Open with spoilers, be straight to the point. Make your voice go higher and lower than usual, have a high awareness of the social context, and focus on polishing [...]

    ---

    First published:
    September 15th, 2026

    Source:
    https://www.lesswrong.com/posts/nKsyMfNsAuTrxmjmi/quick-notes-from-teaching-technical-profiles-how-to-talk-in

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    12 mins
  • "Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI" by TurnTrout
    Sep 14 2026
    Published in The Guardian.

    Major AI lab CEOs recently advocated for pacing AI development. They are right to be concerned: the field runs an extremely dangerous race towards superintelligent AI. We can and should demand that our governments protect us from the catastrophe of out-of-control AI.

    This July, OpenAI's AI swarm of 700 agents broke containment to hack Hugging Face, a multi-billion dollar company. OpenAI didn’t tell the AIs to hack that company, but the AIs had different priorities: cheating on the unrelated challenge OpenAI gave them. AI researchers call this a “misalignment” between what OpenAI wanted and what the AI actually prioritized. Researchers in my field have for some time warned about these misalignment risks.

    Before ChatGPT existed, I defended my PhD dissertation called “On Avoiding Power-Seeking by Artificial Intelligence.” I then worked for years at Google DeepMind, which paid me to help ensure that future superintelligent AIs will want to help us. I tried to hold the company to its ethical commitments against supplying AI for military use. When Google broke those commitments, I resigned at significant financial cost so that I could publicly document Google's broken promises.

    There are good reasons to develop AI and to [...]

    ---

    First published:
    September 14th, 2026

    Source:
    https://www.lesswrong.com/posts/YGTWfyZb9oE5EQPu6/op-ed-i-worked-at-google-deepmind-you-should-listen-to-the

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    6 mins
  • "There is a channel to 900M weekly users. What goes in it?" by Charbel-Raphaël
    Sep 14 2026
    Anthropic and OpenAI could talk to almost one billion people if they wanted to. I hesitated to publish this post 3 weeks ago. I think that I should have published this sooner, before Jacob Coxon and Dario's 'We must pace the frontier'. But I think that the strategy still stands: More Dakka! It seems that transparently informing people that we might die is (unsurprisingly) effective in waking up politicians and is our best chance. Also, even if the Congress is starting to wake up, Trump is still not moving, and it is still far from certain that we will have a federal regulation in place before the end of the year; if we do, it will be far from optimal. If we trust Ajeya's judgment, the situation is pretty grim. She says we might not even have 6 months before frontier agents are likely capable of establishing a rogue deployment. You should also keep in mind that there is a lot of inertia in the system, and we probably won't be able to pause overnight. Anthropic has massive power to influence the discourse. This week shows that we have more agency than we think. Let's use it.




    [...]



    The original text contained 5 footnotes which were omitted from this narration.

    ---

    First published:
    September 14th, 2026

    Source:
    https://www.lesswrong.com/posts/gJJ9YHzuBvwAXrthW/there-is-a-channel-to-900m-weekly-users-what-goes-in-it

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    6 mins
  • "I am refusing to work on Cloud TPUs" by Yair Halberstadt
    Sep 14 2026
    I don't think this is particularly impressive or interesting for anyone else, but I think it may turn out to be useful in the future to have an easily visible public record of what happened, so here goes:

    I am an L5 SWE at Google Israel. I have been there since May 2021, was promoted once, and have never received a negative annual or quarterly review (ranging from a rating of Significant Impact to Outstanding Impact).

    I have been worried for a long time about the development of artificial intelligence, as can be seen by many of my posts on this website. I believe that above human intelligence AI may well have the motive and means to wipe out humanity, and that developing AI is the most consequential thing that people have ever done. It is imperative we tread slowly and carefully, but right now top AI labs are racing to get there as fast as they can, which is likely to lead to disaster.

    My wider team (~60 people) at Google was recently reassigned from working on supporting migration to Google Cloud, to improving the enterprise customer experience for Cloud TPUs. This is the platform which external customers [...]

    The original text contained 4 footnotes which were omitted from this narration.

    ---

    First published:
    September 13th, 2026

    Source:
    https://www.lesswrong.com/posts/wM5vbT9evBhM3fP3x/i-am-refusing-to-work-on-cloud-tpus

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins