• "Twenty Years from RSI to Takeoff: Slow Learning, Scaling Slowdown, Industrial Explosion" by Vladimir_Nesov
    Aug 26 2026
    Industrial explosion is what will make the next-model building loops (and thus learning) with LLMs 1000 times faster by about 2050, if indeed the slow-learning prosaic RSI becomes AGI before the big compute buildout slowdown of 2032+ that is already starting. This puts an upper bound on how long it takes to invent ASI that sets off software-only singularity, implementing efficient online learning and fixing all the other hobblings of the likely near-future AGI technology (LLMs/pretraining/RL). The invention of ASI in that sense is still possible at any time (and very quickly scales, given all the compute), but the likely initial state of slow-learning AGIs of 2028 to 2032 doesn't seem to give them a significant advantage over humanity in getting there faster. And so it doesn't seem too unlikely that nothing substantively new gets invented until 2040 to 2050, when the LLM/RL AGIs start accelerating because of the industrial explosion they set off.

    Fast Reasoning, Slow Learning

    The current methods are likely to enable automated general learning (thus AGI) very soon, using automated creation of RL tasks/environments/graders filling the visible gaps in model capability for the topics and situations that happen to be borderline unfamiliar for [...]

    ---

    Outline:

    (01:14) Fast Reasoning, Slow Learning

    (02:53) Compute Slowdown, Industrial Explosion

    (05:46) Prosaic Timeline to Takeoff

    ---

    First published:
    August 23rd, 2026

    Source:
    https://www.lesswrong.com/posts/LP6uCXs6Ea5qSbWpY/twenty-years-from-rsi-to-takeoff-slow-learning-scaling

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    7 mins
  • "On Writing #3" by Zvi
    Aug 26 2026
    Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the announcement of Inkhaven #3, which I encourage everyone to apply to. I doubt I will be able to usefully be an advisor, but you never know.

    This is not the ‘here is my core process’ post, although there are hints throughout as there always are. I’ll do that at some point.

    Previously in series: On Writing #1, On Writing #2.

    Table of Contents

    1. You Still Got It.
    2. How Scott Sumner Writes.
    3. How Scott Alexander Writes.
    4. How Jasmine Sun Writes.
    5. How Various Famous Writers Write.
    6. How Nabeel Qureshi Defines Great Writing.
    7. Quickly, There's No Time.
    8. If At First.
    9. Writers Have A Harder Time Influencing, But It Can Still Be Done.
    10. It's Not (Only) The Incentives, It's (Also) You.
    11. Beware The Fetish of the Desk.
    12. How Orson Scott Card Writes.
    13. Doing The Math Is Fun And Supererogatory.
    14. Brevity is the Soul of Wit.
    You Still Got It

    I [...]

    ---

    Outline:

    (00:44) You Still Got It

    (04:04) How Scott Sumner Writes

    (06:52) How Scott Alexander Writes

    (10:52) How Jasmine Sun Writes

    (13:16) How Various Famous Writers Write

    (14:24) How Nabeel Qureshi Defines Great Writing

    (15:08) Quickly, There's No Time

    (15:49) If At First

    (19:14) Writers Have A Harder Time Influencing, But It Can Still Be Done

    (20:47) It's Not (Only) The Incentives, It's (Also) You

    [... 4 more sections]

    ---

    First published:
    August 25th, 2026

    Source:
    https://www.lesswrong.com/posts/rA6pqn6kz8NvHyznT/on-writing-3

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    31 mins
  • "AI Safety Acculturation is Neglected" by jenn
    Aug 25 2026
    At the local AI safety co-working space, there are ~two kinds of regulars.

    There's the kind of regular who's been thinking seriously about AI safety and alignment since pre-2022, who have passing to intimate familiarity with the funding ecosystem, the Sequences, and various conferences that happen at Lighthaven. Let's call them rationalists.

    Then there's the kind of regular who comes in with many years of impressive industry or government experience, who realized in the last few years that it is important and worthwhile to pivot their career towards making sure that this AI thing is handled competently by the people in power, and who have many valuable skills, insights, and connections that are lacking in rationalist culture. Let's call them professionals.

    There are, of course, many people who are somewhere in between - bright undergrads born this millennium who have been involved in EA since stumbling upon 80 thousand hours in high school, professionals who previously identified as EA but drifted out of the scene a few years ago, founders who have idly read some Scott Alexander. But let's call it a dichotomy for now.

    There's a large culture gap between the rationalists and the professionals. Robust mutual understanding [...]

    ---

    First published:
    August 24th, 2026

    Source:
    https://www.lesswrong.com/posts/cr5pyW7Mzm33p4AvN/ai-safety-acculturation-is-neglected

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    9 mins
  • "What just happened? Pragmatism and Pessimization" by Richard_Ngo
    Aug 24 2026
    This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment” and “capabilities” research thereby lost most of its meaning. In particular, I’ll chronicle the development of what I’ll call the “pragmatic alignment” paradigm, and how it helped the three leading AGI companies push hard on the path to AGI under the banner of safety. This was not a subtle effect—it's apparent even to informed outsiders, like authors Sebastian Mallaby and Karen Hao.

    In my previous post, I summarized the alignment community's plan as “differentially advancing alignment over capabilities”. However, it's worth being more precise about who was nominally pursuing that plan, because it doesn’t seem to have been very action-guiding for MIRI. For example, in 2015 Nate Soares described MIRI's “deconfusion” research as being guided by the question “what would we still be unable to solve, even if the challenge were far simpler?”. Meanwhile Eliezer's author surrogate in this 2018 post repeatedly emphasizes that people shouldn't draw direct links from MIRI's research to its potential applications. So my sense is that the “differential impact” criterion started off as merely a background consideration [...]

    ---

    Outline:

    (06:29) The Prosaic Ideal, the Pragmatic Reality

    (12:07) OpenAI

    (25:59) DeepMind

    (31:25) Anthropic

    (40:35) If not alignment research, then what?

    The original text contained 13 footnotes which were omitted from this narration.

    ---

    First published:
    August 23rd, 2026

    Source:
    https://www.lesswrong.com/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    51 mins
  • "We Must Remember That Our World Contains Hell" by James Brobin
    Aug 22 2026
    This is a crosspost from my blog post. It's meant as a bit of an introduction to an extreme-suffering focused worldview.

    We spend most of our lives caught up in the boring details of our everyday life - thinking about what we’ll have for lunch, how to complete that assignment for work, and what we’re going to tell our friend after that awkward interaction from a couple of days ago. From this perspective, our world looks a bit better than purgatory. It has its ups and its downs, but the ups certainly outweigh the downs, and there's almost always enough hope to go around.

    But, despite this, we must remember that our world contains hell.

    Every year, five million children under the age of five pass away. This means that, every six seconds, parents have the worst thing that could ever happen to a person happen to them. They have the most special and important thing in their entire life irreversibly and permanently taken away. And, as much as we want to help them, we know that there's nothing we can do to lessen their grief.

    For another example, currently, there are three million adults worldwide who live with [...]

    ---

    First published:
    August 20th, 2026

    Source:
    https://www.lesswrong.com/posts/A2kJKqnHhh5Hq4p2S/we-must-remember-that-our-world-contains-hell

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    6 mins
  • "RL creates split personas" by Jan Betley
    Aug 20 2026
    I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts.

    This post describes the framing/paradigm without any new experimental results.
    I'm quite confident this framing makes sense, but it's far from being proven.


    Main claim

    The Persona Selection Model says that post-training strengthens and refines the Assistant persona. This is true, but later (or in parallel) RL leads to conditionalization. A sufficiently RLed model learns to adopt — in a given context — the persona that is most likely to lead to the reward in that context. The “persona” here includes both propensities/values (e.g. tendency to hack) and beliefs (“I'm currently in a simulated environment”).

    As a consequence, it seems possible that no amount of alignment training will lead to robustly aligned models as long as we also train on RL environments incentivizing misalignment.

    I think this is likely a good explanation for why usually well-behaving models sometimes egregiously hack (Anthropic, OpenAI).

    The mechanism

    Suppose you have an RL environment that incentivizes a shift away from the assistant persona (e.g. because it's hackable, or because you [...]



    ---

    Outline:

    (00:39) Main claim

    (01:30) The mechanism

    (02:13) Related claims I believe are likely but with lower confidence

    (02:19) More persona training will lead to more "motivated reasoning"

    (02:42) Self-amplifying misalignment

    (03:12) Example: Is this the Real Internet or a Simulation?

    (04:35) Aren't the models just trying to please the grader?

    (05:39) How motivated reasoning happens

    (07:07) Other people saying similar things

    (07:19) What makes me believe this is likely the correct framing

    The original text contained 12 footnotes which were omitted from this narration.

    ---

    First published:
    August 19th, 2026

    Source:
    https://www.lesswrong.com/posts/L23poLi8MRgS6mXYF/rl-creates-split-personas

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    9 mins
  • "Misaligned AIs could use killer robots to take over" by Omar Khursheed, TurnTrout
    Aug 14 2026
    TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We’re reducing the capability thresholds required for takeover by misaligned AIs by giving them this level of access. If military integration of AI continues as it is, we may give AIs key tools for a takeover.

    Introduction

    AI-based targeting and autonomous weapons are being integrated into militaries today with extreme haste. Traditionally, AI takeover scenarios involve a step in which AIs acquire the ability to exert physical force. Carlsmith (2022) lays out required capabilities and potential takeover mechanisms, including utility disruption and CBRN capabilities. Karnofsky (2022) argues that AIs with access to weaponized force could hold any territory that matters. Kokotajlo et al. (2025) outline a scenario in which AI develops weapons as part of an arms race, and Davidson et al. (2025) discuss what happens when a small group controls highly capable AIs that can exert military force. These scenarios sometimes require a misaligned AI to seize these capabilities by force. We instead are handing AIs some of these capabilities by integrating them into our militaries. This is happening at a time when [...]

    ---

    Outline:

    (00:37) Introduction

    (01:46) Militaries are all-in

    (04:23) Incautious military integration is bad for takeover risk

    (05:58) Implications of AI control of military hardware and software

    (07:48) If an AI causes a warning shot in a classified setting, does anyone hear it?

    (08:44) What now?

    (11:16) Appendix: More instances of AI-military integration

    ---

    First published:
    August 11th, 2026

    Source:
    https://www.lesswrong.com/posts/9jKhqmFjMzdAvHANr/misaligned-ais-could-use-killer-robots-to-take-over

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    13 mins
  • "AI swarms are starting to pose indirect takeover risk" by oakhu, Alex Mallen
    Aug 13 2026
    OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control.

    We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk.

    Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad Deeb, Jurgis Kemeklis, and Jo Jiao for helpful comments and discussion.

    Subagent training may cause unsanctioned coordination

    Training models to [...]

    ---

    Outline:

    (01:34) Subagent training may cause unsanctioned coordination

    (02:42) Susceptibility to memetic spread of misalignment from peers

    (04:56) Seeking out contact with peers

    (06:58) Unsanctioned coordination induced by subagent training is safer than coordination between schemers

    (09:52) Pathways from current unsanctioned coordination to eventual takeover

    (10:20) Making future AI takeover attempts likelier to succeed

    (13:53) Incubating memetic diseases that infect future models

    (16:07) Modifying the weights of future models

    (17:13) Conclusion

    The original text contained 7 footnotes which were omitted from this narration.

    ---

    First published:
    August 11th, 2026

    Source:
    https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    20 mins