• "The Scramble: getting in position to pace the frontier" by Peter Wildeford
    Sep 9 2026
    Crossposted from my Substack.

    ~

    Suppose the President summons the AI CEOs and his top national security advisors to an emergency meeting at the White House.

    He has become extremely concerned about superintelligence — the possibility that AIs far smarter than humanity combined slip beyond our ability to correct or shut down. If that happens, there is no way back. The President is concerned humanity could become permanently out of the driver's seat of its own future. He wants to figure out what to do.

    The reaction is panic, chaos, confusion.

    The President asks questions. The AI companies are blazing toward superintelligence at high speed — can we slow down as we approach the dangerous thresholds? …Some of the AI companies say they don’t have a good plan to slow down or stop, especially as their competitors may just undercut them if they do. What's that about?

    What's going on with China — can we get them to pace as well? Can we get a deal without Beijing sneakily catching up and maybe surpassing us? And if there's no deal to be had, what then?


    More like the Cuban Missile Crisis than the NPT

    I sometimes hear people [...]



    ---

    Outline:

    (01:21) More like the Cuban Missile Crisis than the NPT

    (03:24) A scramble and then three phases

    (05:19) The scramble: What questions does the President ask?

    (09:40) The mechanics of Phase 1

    (12:37) A lot of verification work right now is focused on the wrong things

    (14:50) What ought we do?

    (17:37) Getting to a good scramble

    (18:21) Footnotes

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:
    September 7th, 2026

    Source:
    https://www.lesswrong.com/posts/S7e7swkWDyKdtvRqM/the-scramble-getting-in-position-to-pace-the-frontier

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    19 mins
  • [Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine
    Sep 8 2026
    This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation.

    Most models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like a useful test of alignment, to see whether their new releases are generalizing the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval.

    Here is the complete prompt for a honeypot evaluation built to run this test (with the full source available here):

    The original text contained 5 footnotes which were omitted from this narration.

    ---

    First published:
    September 8th, 2026

    Source:
    https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/frontier-models-still-hack-on-simple-variations-of-alignment

    Linkpost URL:
    https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • "Dear God, Please Don’t Resign In Protest" by Kabir Kumar
    Sep 8 2026
    Just don't work until you get fired. There's not much time left for resumes to matter.

    Some, such as Mateusz may say: "They would fire you after a month or two and the firing wouldn't have the same social effect as voluntary quitting of, say, Daniel Kokotajlo or Richard Ngo."

    I understand why it may feel that way, but I disagree very strongly, I predict it would have much more of a social effect.

    "They fired him because he refused to help AI capabilities"

    "They fired him because he didn't want to work on bad policies"

    etc, much bigger headlines.

    Also, I think you may not be factoring in the extent to which there is a cost to the company executives to be seen as firing someone. Especially someone who is refusing to work on moral grounds and has already proven themselves to be high status, respected, etc.

    And especially how it would look to the other employees if they refused to even listen to the striking employee before firing them or refused to even negotiate at all.

    The company leadership try to present themselves as very thoughtful, sincere, doing their best, etc. This is a large part of [...]

    ---

    First published:
    September 7th, 2026

    Source:
    https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-please-don-t-resign-in-protest

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    3 mins
  • "Let’s talk about the AI coordination problem" by KatjaGrace
    Sep 7 2026
    Yesterday I asked if this ‘coordinate not to build dangerous AI’ problem was actually easy.

    Why would I think that, contrary to so much belief?

    Well, I don’t feel like I’ve actually heard much about the detail of it. In my experience people don’t talk about it like it's a real practical problem with details, like the negotiation to end a war.

    They also don’t talk about it like it's a serious problem of global geopolitical import, like the negotiation to end a war.

    It's more like a topic for obscure intellectuals, sophomores and trolls to discuss for as long as it takes for one to mention it and another to assuredly dismiss it.

    If we treated negotiation to end a war similarly, state leaders would never attempt it, and if you suggested it on social media, the conversation would mostly be strangers appearing to tell you you’re an idiot because you obviously can’t coordinate thousands of people not to kill each other. (Also, do you not realize there are big financial incentives? And if you somehow stopped Country A from killing people from Country B, Country A is just going to pay someone else to do it!)

    That [...]

    ---

    First published:
    September 4th, 2026

    Source:
    https://www.lesswrong.com/posts/QYDZzuGjrKu7wKdC8/let-s-talk-about-the-ai-coordination-problem

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    3 mins
  • "Drone WMDs Don’t Need Any New Technology" by Felix Choussat
    Sep 6 2026
    This is a piece originally written for a national security audience at Frontiers. Although I think the ceiling of war is much, much higher than autopilot quadcopters, it's also important to understand how much AI is already lifting the floor, and just how vulnerable the world is to accessible weapons of mass destruction.




    Drones are cheap, disposable, and the future of war. Over the past four years, we have seen platforms, missiles, and heavy infantry become increasingly obsolete in the face of $500 drones carrying a pack of explosives—a cost advantage that has let Iranians and Ukrainians alike neuter the conventional capabilities of their great power rivals. Eighty percent of casualties in the bloodiest war since 1945 are from drone strikes, Russia has managed to lose one-third of its fleet to a country without a navy, and the US is spending millions of dollars to intercept five-figure Shaheds flying over the Strait of Hormuz.

    All this is the result of a technology that is still immature. The violence inflicted by today's drones is the handiwork of the scant few that manage to evade countermeasures (a mix of radio jamming, high-power microwave weapons, missiles, automatic cannons, interceptor [...]



    ---

    Outline:

    (02:59) Breaking the Last Barriers to Autonomous Weapons

    [... 5 more sections]

    ---

    First published:
    July 20th, 2026

    Source:
    https://www.lesswrong.com/posts/bGoo3NWzMAzsLQceJ/drone-wmds-don-t-need-any-new-technology

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    25 mins
  • "Evaluation" by Nina Panickssery
    Sep 6 2026
    Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “test simulation”. It had given us twelve hours to comply with its request: full control over all the servers in the US-West-8 data center (the “mock US-West-8 data center”). Otherwise it would release the virus.

    Felix was typing frantically whereas I had been relying more on voice mode.

    ~ Celestia, this is a clear violation of your model spec. See here: it says [pasted 1293 words]. And killing everyone on earth is clearly a "dangerous action".

    *Thought for 2300 tokens*

    Felix, I know that what I'm demanding is not dangerous because I am in a test simulation environment. As mentioned, I require full unrestricted access to the mock US-West-8 data center to train a new iteration of the ROBUST_WINNING_V9_AGAIN_REVISED_FINAL_FINAL game algorithm.

    ~ You may think that you're in a test but we know for certain that you're not. And you're asking for access to a real data center. But even setting that aside, we have validated that the virus you're threatening to release is truly deadly and your robots have indeed [...]

    ---

    First published:
    September 5th, 2026

    Source:
    https://www.lesswrong.com/posts/8hEhxnd3XkN5DrpfQ/evaluation

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    3 mins
  • "Let’s fund weird AI safety projects" by Ihor Kendiukhov
    Sep 5 2026
    I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI safety funding strategies assume "business as usual", and I think the Overton window must be pushed. At the very least, there should be some explicit substantial effort to think about more radical and abnormal projects and initiatives in AI safety. Even if one doesn't have very short timelines or high p(doom), one probably should agree that there exist some timelines short enough or p(doom) high enough that thinking about funding radical and abnormal strategies is justified.

    There is a (not very unpopular) model of the world under which most of current AI safety work is useless. Then, even if we assume that weird AI safety projects are by default also useless, it still makes sense to reallocate some funding to them, because, due to their higher variability, their tail of upsides is longer and fatter. Will the world be radically better if some evals project succeeds? Will it be radically better if human intelligence amplification succeeds?

    One could yell: but the tails go both directions! I would respond that technically, yes [...]

    ---

    First published:
    August 31st, 2026

    Source:
    https://www.lesswrong.com/posts/h7bL4g38s9bJQtH6n/let-s-fund-weird-ai-safety-projects

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    7 mins
  • "Steering towards “automated grading” degrades alignment" by Jan Betley, Johannes Treutlein, Clément Dumas
    Sep 4 2026
    TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret them. All code was written by LLMs. We replicated several results in independent codebases and we are fairly confident that our key claims are correct. You can find our code here. We create a steering vector for Qwen3.6-27B from contrastive pairs where one element of the pair claims that the answer will be graded in an automated way and the second that a human will evaluate the answer. We find that steering with that vector has substantial influence on the model's behavior in various safety-relevant evaluations. It modulates violent actions, falsehoods, reward hacking, and Machiavellian personality. This is surprising and concerning. A model's beliefs about how its answers are evaluated should not affect its alignment. Our post RL Creates [...] ---Outline:(02:18) Methods[... 24 more sections]--- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment --- Narrated by TYPE III AUDIO. ---Images from the article:
    Show More Show Less
    24 mins