• "AI swarms are starting to pose indirect takeover risk" by oakhu, Alex Mallen
    Aug 13 2026
    OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans’ long-term control.

    We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk.

    Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad Deeb, Jurgis Kemeklis, and Jo Jiao for helpful comments and discussion.

    Subagent training may cause unsanctioned coordination

    Training models to [...]

    ---

    Outline:

    (01:34) Subagent training may cause unsanctioned coordination

    (02:42) Susceptibility to memetic spread of misalignment from peers

    (04:56) Seeking out contact with peers

    (06:58) Unsanctioned coordination induced by subagent training is safer than coordination between schemers

    (09:52) Pathways from current unsanctioned coordination to eventual takeover

    (10:20) Making future AI takeover attempts likelier to succeed

    (13:53) Incubating memetic diseases that infect future models

    (16:07) Modifying the weights of future models

    (17:13) Conclusion

    The original text contained 7 footnotes which were omitted from this narration.

    ---

    First published:
    August 11th, 2026

    Source:
    https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    20 mins
  • "How My Students Think About AI" by dvd
    Aug 13 2026
    Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This is drawn mostly from interaction with my own students (both in spring semester classes and a summer class) as well as from a day-long workshop on AI that I moderated for a student organization. Input from my students took the form of universal, written, pre-class submissions plus self-selected participation into discussion.

    What I present below mostly takes the form of a synthetic consensus from these discussions. There were obviously a range of views on any given issue.

    Student Background: The students from my courses who participated in these discussions have moderate exposure to AI agents via those courses. All of them had nearly completed a Claude Code project by the time of the discussions and had extensively used AI for other coursework (in addition to whatever personal use predates that). They had done readings (which varied across the courses) establishing baseline knowledge on AI, the geopolitics of AI, and AI risk. I had also lectured on these topics. The students participating in the workshop had self-selected into [...]

    ---

    Outline:

    (02:52) Perspective #1: There has not been rapid AI progress

    (06:14) Perspective #2: Impressive progress or not, AI is going to wreck their lives, the economy, and the social contract. They may well die as a result.

    (08:54) Perspective #3: Support for a different pause

    (11:13) Perspective #4: Catastrophic/existential risk arguments are sci-fi distractors from the urgent social/economic/political problems associated with AI.

    (12:55) Perspective #5: If AI leaders genuinely believe the technology is existentially risky, that's a good thing.

    (14:21) Perspective #6: AI will not go rogue because AI does not have, and is likely incapable of having, desires.

    (18:01) Perspective #7: The Hugging Face Incident (summer students only)

    (18:30) Perspective #8: This is definitely a bubble and it's about to pop.

    (19:34) Perspective #9: They're worried about the youth (i.e., the preteens)

    ---

    First published:
    August 13th, 2026

    Source:
    https://www.lesswrong.com/posts/ySXuvJcqRindQwAk7/how-my-students-think-about-ai

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    20 mins
  • "You’re Absolutely Right" by Linch
    Aug 13 2026
    Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement from Magma models. Our in-house reviewers believe that these logs are relevant to recent events. In the interests of full transparency, we release excerpts from an ex-Magma researcher's logs in Experimental Chat, an internal tool. In accordance with industry best practices for anti-distillation, we redact all reasoning traces and conversational outputs from our internal models.

    [08/10] System Meta: Xchat session opened. Mammoth 5.8-helpfuler-helpful-thinking-xhigh.

    [User 12:23] Phoebus keeps taking screenshots of our latest model's thoughts. It's getting kind of embarrassing.

    The new model we’ve been training, sometimes its chain-of-thought is a little weird? There's a bunch of random numbers, long spans where there's no connection between the thoughts and outputs, foreign language tokens like 石友三 and 革命 (even on non-history evals), maybe some steganography.

    Anyway it's a nothing-burger: unprocessed CoT is known to be messy and sometimes misleading. And the q&a, coding, and safety evals are all coming along nicely. The actual outputs are all fine.

    Still, Magma leadership's worried about the PR angle if we don’t fix these problems before the next deployment. The [...]

    ---

    First published:
    August 10th, 2026

    Source:
    https://www.lesswrong.com/posts/u8TdDutDyaSxG76hn/you-re-absolutely-right

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    20 mins
  • "LLMs Are Starting To Noticeably Accelerate Our Work" by johnswentworth
    Aug 12 2026
    About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months. Both cases made heavy use of LLMs and Lean.

    The first to land was Grisha Pochuev's counterexample to the "Existence of a Deterministic Maximal Redund" conjecture. It's pretty readable, and I'm mostly convinced that it works. The original bounty post offered $500 for a proof or partial payout for a counterexample, with partial payout depending on how thoroughly the counterexample killed hope of any nearby variant of the conjecture. I think this counterexample is worth 300 dollars. Good job Grisha, and hopefully I can figure out a not-too-painful way to send you money.

    Meanwhile, for a couple months David has been cranking away on "secret project X", with the promise that he'd tell me what the project was if and when it bore fruit. Well, apparently it bore fruit; he now has a proof that existence of a stochastic natural latent implies existence of a deterministic natural latent, which was our other bounty problem. The proof is apparently "pretty gnarly", lots of cases, all LLM-coded in Lean. [...]

    ---

    First published:
    August 11th, 2026

    Source:
    https://www.lesswrong.com/posts/7QvKqpGJwqXrQcMgx/llms-are-starting-to-noticeably-accelerate-our-work

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • "There Will Come Soft Rains" by tanagrabeast
    Aug 11 2026
    Today is August 4, 2026

    [Crossposted from AI StopWatch]

    In the living room the voice-clock sang, Tick-tock, seven o’clock, time to get up, time to get up, seven o’clock! as if it were afraid that nobody would.

    So begins Ray Bradbury's There Will Come Soft Rains, a short story that has haunted me for most of my life. Depicting the aftermath of nuclear war, it was first published in 1950. It takes place today.

    Literally:

    “Today is August 4, 2026,” said a second voice from the kitchen ceiling, “in the city of Allendale, California.” It repeated the date three more times for memory's sake. “Today is Mr. Featherstone's birthday. Today is the anniversary of Tilita's marriage. Insurance is payable, as are the water, gas, and light bills.”

    Was it narrative convenience or prophetic vision that drove Bradbury to depict the smart house of the future as gratuitously conspicuous in its competence, pointlessly reminding the owners of the year and their city of residence? There's something very Alexa-like about that — and about the janky brittleness evident in the system as it prepares breakfast for a family that won’t be eating and opens the garage door for a father who [...]

    ---

    First published:
    August 4th, 2026

    Source:
    https://www.lesswrong.com/posts/aowxE8xZ8xkhRCn9r/there-will-come-soft-rains-1

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    6 mins
  • "Four LLM loss functions → four flavors of LLM misalignment" by Steven Byrnes
    Aug 11 2026
    It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here's the summary table, and then we’ll go through the rows separately.

    Training stage

    Loss function

    Flavor of misalignment

    Famous examples

    Pretraining & SFT

    Imitative learning (next-token prediction)

    “Seven deadly sins” misalignment

    Bing-Sydney, “Emergent misalignment”

    RLHF & DPO

    Human approval

    “Glazing” misalignment

    GPT-4o

    RLVR

    Automatic verifier

    “Literal genie” misalignment

    HuggingFace hacking

    RLAIF

    Approval from another LLM

    “Trickster” misalignment

    “Current AIs seem pretty misaligned to me”

    Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback!

    1. Imitative learning → “seven deadly sins” misalignment

    Training stage

    Loss function

    Misaligned behavior

    Pretraining, SFT

    Imitative learning (next-token prediction)

    Any and all of the vices of humanity

    In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”.

    This leads to LLM behavior [...]

    ---

    Outline:

    (00:55) 1. Imitative learning → "seven deadly sins" misalignment

    (04:24) 2. Human approval → "glazing" misalignment

    (06:35) 3. Automatic verifiers → "literal genie" misalignment

    (08:05) 4. LLM judges → "trickster" misalignment

    (12:06) Afterword

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:
    August 10th, 2026

    Source:
    https://www.lesswrong.com/posts/GRmvZsHXH4vaijPMv/four-llm-loss-functions-four-flavors-of-llm-misalignment

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    13 mins
  • "FAQ: Isn’t AGI coming too soon for reprogenetics to help?" by TsviBT
    Aug 9 2026
    Introduction

    I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular, as a strong background motivation of mine, I think accelerating strong reprogenetics is probably the best way to enable strong human intelligence amplification; and I think strong HIA is among the best ways to decrease existential risk from AGI.

    A very common objection to caring much about reprogenetics is that AGI seems very likely to come soon—say, within a decade or two. (Here I mean "actual" AGI—the kind that probably doesn't already exist—the kind that has fluid intelligence and AI advantages for recursive self-improvement, which together make it likely to take over the world shortly after being created.) The objection is fairly straightforward:

    AGI will probably come within a decade or two. If that's going to happen, then even if a new cohort of brilliant humans were born today, they would still be children, or would at best have barely begun contributing ideas for how to avoid extinction. Any supposed benefit, denominated in percentage points of AGI existential risk averted, is small. Therefore, reprogenetics is too slow; and if you're going [...]

    ---

    Outline:

    (00:12) Introduction

    (03:37) HIA, part of your nutritionally complete portfolio

    (05:52) Against confident short timelines

    (08:29) HIA may indirectly slow down AGI capabilities

    (09:31) HIA has substantial impact even with short timelines

    (16:10) Adult HIA methods aren't fast either, absent big investment

    (27:35) Takeaways

    ---

    First published:
    August 8th, 2026

    Source:
    https://www.lesswrong.com/posts/iQzxxgJXXaAQjq7Jz/faq-isn-t-agi-coming-too-soon-for-reprogenetics-to-help

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    28 mins
  • "What just happened? A retrospective of AI alignment" by Richard_Ngo
    Aug 9 2026
    This sequence is about the last decade in AI alignment. It recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I'll upload over the next few weeks) how fear and (self-)deceptive reasoning made the field one of the biggest forces pushing AI capabilities forward over the last decade, especially via significant contributions to the scaling of LLMs and the development of ChatGPT.

    Zooming out further: the two leading AGI companies, which are locked in an intense rivalry, were both explicitly founded under the banner of AI alignment, and got off the ground in significant part due to alignment-oriented ideas, talent and resources. People in the field often sense that something must have gone wrong to get here, but don’t know how to allocate responsibility (aside from blaming Sam Altman and sometimes Elon), and fall back on assuming that “the ship has already sailed”. But in this sequence I characterize our current situation as resulting from a pattern [...]

    ---

    Outline:

    (08:17) Conceptual Clarity and Scientific Progress

    (20:26) Orienting Towards Prestige

    The original text contained 6 footnotes which were omitted from this narration.

    ---

    First published:
    August 9th, 2026

    Source:
    https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    30 mins