• "Jensen Huang Says If We Cannot Align AI, Shut Down the AI Labs" by Ben Pace
    Sep 24 2026
    I was very surprised today on a podcast to hear Jensen Huang plainly state that if they cannot align the AIs, then the labs must shut down.

    The context I have on Huang is that he has run NVIDIA for 30+ years, which has become the most valuable company in the world due to the AI boom. My understanding is that he has repeatedly encouraged the US President (with whom he is on friendly terms) to continue to support AI, and dismissed AI talk as "sci-fi".

    If you haven't seen, his biographer has incredible quotes of him being pressed on risks from AI, where Jensen gets furious.

    “This cannot be a ridiculous sci-fi story,” he said. He gestured to his frozen PR reps at the end of the table. “Do you guys understand? I didn’t grow up on a bunch of sci-fi stories, and this is not a sci-fi movie. These are serious people doing serious work!” he said. “This is not a freaking joke! This is not a repeat of Arthur C. Clarke. I didn’t read his fucking books. I don’t care about those books! It's not– we’re not a sci-fi repeat! This company is not a [...]

    ---

    First published:
    September 23rd, 2026

    Source:
    https://www.lesswrong.com/posts/cmdbNijFsopqfqEq7/jensen-huang-says-if-we-cannot-align-ai-shut-down-the-ai

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    5 mins
  • "Alignment Midtraining Cracks Under Pressure" by J Bostock, sidbaines, Daniel Tan, draganover, ma-rmartinez
    Sep 23 2026
    TL;DR

    We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.

    For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190 million tokens of midtrained motivations are overpowered by a relatively tiny amount (~50 thousand tokens) of competing finetuning data. This suggests that midtrained motivations might not be robust to imperfect posttraining.

    Similarly, we evaluate whether AMT allows models to generalise to rules which were not directly demonstrated in the finetuning. We find that the capacity for such generalisation is surprisingly low. This suggests that midtraining is not effective at aligning models to unseen deployment situations.

    In one experiment, we midtrained GLM-4.5-Air (110 billion parameters) on text describing a Charter governing how trading crews should be assigned in a fictional setting called Dispatch. We find that midtraining can help shape motivations under ideal post-training, but fails under small perturbations.

    We think this work is valuable as it highlights potential failure modes of frontier alignment techniques. We encourage others to do more red-teaming of labs' alignment [...]

    ---

    Outline:

    (00:12) TL;DR

    [... 7 more sections]

    ---

    First published:
    September 21st, 2026

    Source:
    https://www.lesswrong.com/posts/QH86EzNsjRw3wtCGs/alignment-midtraining-cracks-under-pressure

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    15 mins
  • "Swarm Scaling" by Toby_Ord
    Sep 22 2026
    Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm?

    We’ve seen two large and extremely capable swarms from OpenAI in the last few months:

    • 1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched a sophisticated criminal attack on the AI company Hugging Face.
    • A swarm of 10,000 agents solved a version of the longstanding Navier-Stokes problem in mathematics. It took them just 88 hours to do so, in which time they sent 5 million messages to each other and used 300 billion tokens.
    No doubt we will soon see even larger swarms with even more impressive capabilities. But they are not cheap. It is estimated that the swarm of 10,000 agents cost about 20 million dollars at API prices. So while they are very powerful, it will be some time before we see the million-fold reduction in cost needed for this level of power to [...]

    ---

    Outline:

    (02:14) HOW DO SWARMS SCALE?

    [... 2 more sections]

    ---

    First published:
    September 21st, 2026

    Source:
    https://www.lesswrong.com/posts/6cb7qd3RSkgnviCpf/swarm-scaling

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    16 mins
  • "We’ve saved the world before: what the ozone hole teaches us about AI" by leogao
    Sep 22 2026
    It might destroy the world, despite passing every known safety test. If we wait for a “warning shot” before we act, it might be too late. And action requires global coordination, because if anyone makes it, everyone dies. Sound familiar?

    It should, because it already happened half a century ago, with chlorofluorocarbons (CFCs). Despite seemingly impossible odds, we got our act together and completely solved the problem through unprecedentedly successful international coordination. The Montreal Protocol banning CFCs, signed 39 years ago today, is the only treaty that has ever been ratified by every single country in the entire world.

    Total Montreal protocol victory

    Making AI go well is going to be a lot harder than fixing the ozone hole. Nonetheless, the similarity is uncanny, and we don’t have any other choice. Understanding how we did the impossible once before may teach us something about how to do it again.

    The theory is born

    The year is 1973. The slow televised unraveling of the Nixon administration is already well underway. DDT finally got banned last year by the newly created EPA. A river got so polluted that it literally caught on fire.

    The Cuyahoga River Fire

    Environmentalism looms large in [...]

    ---

    Outline:

    (01:24) The theory is born

    [... 7 more sections]

    ---

    First published:
    September 20th, 2026

    Source:
    https://www.lesswrong.com/posts/zxXPEtSSSEdwpjopb/we-ve-saved-the-world-before-what-the-ozone-hole-teaches-us

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    18 mins
  • "The Anatomy of a Chinese AI Researcher" by CMLKevin
    Sep 21 2026
    The Chinese AI researcher has read the Three Body Problem series of sci-fi novels since high school, and understand the concept of existential risk vaguely.

    He is fascinated by Ye Wenjie, the researcher that turned against humanity in that book, and decides that in the future if AI progress leads to a superior intelligence, he might be tempted to become Ye if there's no good alternative.

    He performs the duties of capabilities research in a Chinese frontier lab, seeking to one day achieve parity with Western companies, though he knows this is difficult. He has a mentality of hillclimbing, believing that the progress of a future technology is highly uncertain and even unknowable, and so him and his peers could only tread one step at a time.

    He looks at the western world and sees what is typical when a great technology is developed: the first mover will decide to impose restrictions to further their lead, while latecomers should use whatever means necessary to widen access to the whole world. He thinks of the AI chip restrictions as evidence of this.

    He uses Anthropic and OpenAI models regularly in his day to day work. He [...]



    ---

    First published:
    September 19th, 2026

    Source:
    https://www.lesswrong.com/posts/qmxkHm2dTLKG6GZ6i/the-anatomy-of-a-chinese-ai-researcher

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    3 mins
  • "Common mistakes in AI safety group organizing" by Nikola Jurkovic
    Sep 21 2026
    Back in the day, I was a very active AI safety group organizer. I commonly notice people making the same mistakes across many clubs. I have written down a list of some of these mistakes hoping people will avoid them in the future:

    • Reading groups often require that people read things before meetings. This is a mistake. People often don't do the readings. And the lack of common knowledge that everyone has read the reading degrades the conversation quality.
      • Instead, have longer meetings, serve food (so, lunch/dinner meeting slots), and read during the actual meeting.
    • Reading groups often don't sort people into cohorts properly. Mainly, they fail at clustering people into clusters of roughly equal ML knowledge and age. Grad students don't want to discuss a paper with freshmen. People with lots of ML knowledge don't want to discuss a paper with people with no ML knowledge.
      • Instead, group people with people similar to them in ML knowledge and age.
    • Reading groups often rely on digital materials instead of physical printouts. Screens are distracting and there is no common knowledge that people are paying attention.
      • Neatly print every reading ahead of time instead.
    • Clubs [...]
    ---

    First published:
    September 19th, 2026

    Source:
    https://www.lesswrong.com/posts/XFzqDJjAJBt8fkn8i/common-mistakes-in-ai-safety-group-organizing

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • "Please Give Them a Chance: On China, Rationalism, and AI Safety" by gzjw
    Sep 21 2026
    When I finished HPMOR, I immediately knew it was the best novel I had read in more than a decade. I only wished I had found it sooner.




    When I started reading The Sequences, I discovered that the Chinese translation group had translated only the first volume. When I graduated from university, two years ago, AI translation had only just become good enough to convey the meaning of an article with reasonable accuracy. It was only about a year and a half ago that I truly found my way here and began engaging seriously with rationalism.




    My score on the Chinese college entrance exam was only slightly above the cutoff for what was then called a first-tier university. At university, my grades were near the bottom of my year, and I almost failed to graduate. It is probably fair to say that the vast majority of graduates from first-tier Chinese universities are smarter and more capable than I am.




    English has always been my worst subject. From childhood through school, I could barely pass it.




    I have now been working for two and a half years and have saved about 15,000 [...]









    ---

    First published:
    September 20th, 2026

    Source:
    https://www.lesswrong.com/posts/GoX3uYQ4QN5HKvL7u/please-give-them-a-chance-on-china-rationalism-and-ai-safety

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    9 mins
  • "Why I Stay Off Twitter" by jefftk
    Sep 20 2026
    I avoid Twitter (𝕏) for similar reasons to drugs: I think it would change me for the worse, and I would be unable to give it up.

    After staying off Twitter reasonably successfully for years, I cross-posted my AI Tweets there a few weeks ago. I had something very Twitter-shaped to say, and I thought it was important to get out, so I do think this was worth it. And it all went well: none of this is complaining about the comments I got there.

    Coming back a few times to check notifications, however, it's been very good at baiting me: Tweets that are confidently wrong in cases where I have relevant and uncommon knowledge. The pull to dive in and share what I know is very strong! Then this bleeds over to the far broader case where people are wrong, and you have a large potential time sink.

    If it were just the time sink, I'd stop resisting. I spend some time on HN and Reddit, and to the extent that Twitter could substitute for that by showing me things I was more interested in, that wouldn't be an issue. The real problem [...]

    ---

    First published:
    September 19th, 2026

    Source:
    https://www.lesswrong.com/posts/tvwtwgcujTfep4HgY/why-i-stay-off-twitter

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins