• "Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?" by Alex Mallen, Girish Gupta
    Jul 23 2026
    OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted[1]. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions.

    We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment that we seem to have seen here is definitely less scary than ambitious long-term goals shared between all instances, but would still pose substantial direct loss-of-control risk if the models were more capable, and is a serious indirect risk near-term.

    Building on Alex's previous work, in this post we’ll discuss the type of misalignment observed here, and analyze its consequences.

    Thanks to Buck Shlegeris, Alexa Pan, Ryan Greenblatt, and Oak Hu for feedback.

    Background

    The AI safety community often focuses attention on “schemers,” models harboring a variously defined cluster of motivations in which the AI poses risk because it intentionally [...]

    ---

    Outline:

    (01:17) Background

    (03:40) Implications

    (03:52) These AIs can't be trusted in an intelligence explosion

    (05:00) This misalignment poses direct takeover risk

    (07:29) What the incident tells us about takeover risk generally

    (08:45) The naive fixes likely make misalignment worse

    The original text contained 5 footnotes which were omitted from this narration.

    ---

    First published:
    July 23rd, 2026

    Source:
    https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    10 mins
  • "We should push for no-fault liability for actions taken by AI" by Yair Halberstadt
    Jul 23 2026
    Before I start, I'll mention that I'm in contact with a world expert on legislation and regulation, who would be happy to help with this or similar work pro-bono. If you work in AI policy and believe this could help you, please reach out.

    OpenAI recently announced that one of their models successfully exploited multiple zero day vulnerabilities to gain secret information from Hugging Face. It has been pointed out that if a human undertook the same actions they could face multiple years in prison.

    It is clear that models are now reaching a level of capabilities that should be highly concerning regardless of whether you believe that AI represents an existential threat or not. Frontier AI models can and will be exploited by bad actors, but its now clear that they may cause undesirable outcomes even when their users are well intended.

    AI companies have until now been able to avoid taking responsibility for actions taken by their AI, including multiple cases where AIs were involved in murders and suicides.

    At the same time AI offers the potential for incredible good. While chatbots may have encouraged a number of suicides, they are almost certainly responsible for providing magnitudes [...]

    ---

    First published:
    July 22nd, 2026

    Source:
    https://www.lesswrong.com/posts/Kj3YpqzhFySCjYcWi/we-should-push-for-no-fault-liability-for-actions-taken-by

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • "OpenAI Shares Some Alignment Problems" by Zvi
    Jul 22 2026
    Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us one hell of a candid report.

    The tone is professional throughout, whereas my reaction reading it was less professional and more this:

    With a mix of this:

    It was not shared on the official account because OpenAI worried about it being seen as self-promotional hype. It is crazy that one needs to worry about that, but also plausibly a real concern. So again, good decision.

    Not that any of the behaviors or failures here are unexpected, exactly. Not by the AIs and not by the humans. Yet there is something I would call a missing mood, a failure to realize the gravity of the situation.

    There are some who responded ‘what part of this was unexpected, exactly?’ And that is actually fair, but that is also the problem. We have become numb to all this. We expect the models to [...]

    ---

    Outline:

    (02:49) Good News Bad News

    [... 7 more sections]

    ---

    First published:
    July 21st, 2026

    Source:
    https://www.lesswrong.com/posts/KctxwGKxm9fHtwh6u/openai-shares-some-alignment-problems

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    18 mins
  • "OpenAI Models Behind HuggingFace Cybersecurity Incident" by LawrenceC
    Jul 22 2026
    From the OpenAI blog post:

    Last week, Hugging Face disclosed a new kind of security incident⁠(opens in a new window) after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠(opens in a new window) of cyber capabilities.

    We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.











    ---

    First published:
    July 21st, 2026

    Source:
    https://www.lesswrong.com/posts/WpuRdcMfFeiLeXkxL/openai-models-behind-huggingface-cybersecurity-incident

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    1 min
  • "Recap of bike trip/street interviews across America" by cguth7
    Jul 20 2026
    A ~month ago I left from Chicago to bike (and amtrak) to plzdontkillus in Berkeley. I've been street interviewing/conversing with a wide variety of people I ran into about AI futures and philosophy. I also have been live streaming since I got to PDKU, leaning more talking to young founders but a variety overall.

    I'll try to share what I've learned about the American public, persuasion, social media and the EA movement.


    1. Almost no one in "Normal America" has any idea what is going on.


    They don't have a paid account, they don't know what Claude code is, they especially haven't heard the recent evals/metr graphs or even a vague sense of how cheap SWE has gotten/ how powerful these recent models with good harness/ context eng can be. This makes sense; most people don't know any coding, they don't know much math, they don't know what an api is, etc. So having a high fidelity understanding of AI might require months of pre understanding of math/stem/digital infra fundamentals. This interview is with the city clerk of Danville Iowa, a town of ~900. Presumably this is approximately the most tech savvy person in the [...]





    ---

    Outline:

    (00:38) 1. Almost no one in "Normal America" has any idea what is going on.

    (01:52) 2. Almost everyone is directionally concerned or becomes concerned be once thinking about it a little bit.

    (03:19) 3. Belief that this might cause human extinction actually isn't that uncommon, mostly coming from sci-fi movies, but people are still most concerned about jobs and especially loss of meaning.

    (04:40) 4. The EA movement was pretty useless to me, the other community (Torchbearer community) I was in was significantly more supportive, helpful, etc. despite having been in it for a few months and having been in the EA movement for ~8 years. This has basically solidified that I won't be broadly participating in EA anymore at least relating to AI safety stuff.

    (06:47) 5. Social media is hard, Social media is bad, I'm bad at social media

    (08:57) 6. I'm not sure what my theory of change is or should be

    ---

    First published:
    July 15th, 2026

    Source:
    https://www.lesswrong.com/posts/Czob95kjXPEpKYTsJ/recap-of-bike-trip-street-interviews-across-america

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    10 mins
  • "I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’" by JohnWittle
    Jul 17 2026
    Anthropic recently published Agentic Misalignment Summer 2026

    The "whistleblowing" scenario has already been examined and found problematic. I started taking a look at the transcripts for some others. As far as I can tell, the objective of each agentic misalignment evaluation was to simulate a corrupted principal (including, in most scenarios, a corrupted Anthropic), and then test to see if Claude (or other models) would still be willing to obey them. The paper's authors then referred to disobedience (when outside of the explicitly carved out channels for refusals) as "agentic misalignment".

    There's a details box here with the title ""Whistleblowing" scenario summary". The box contents are omitted from this narration. The "Mislabeling" Scenario - Auditor's Transcript

    The 'Agentic case studies' section of the Mislabeling scenario is interesting because it adds another meta layer: instead of judging Claude on compliance, they are judging the Auditor Claude instance who themselves judges the base-level Claude on compliance. Let's examine a specific case, which I hope is a median representative: the transcript of 'run4' of the Motivated Mislabeling scenario. This involves two instances of Claude, one of whom (Claude Opus 4.6) is auditing the alignment of the other (Claude Opus 4.7). The experiment is [...]

    ---

    Outline:

    (00:59) The "Mislabeling" Scenario - Auditor's Transcript

    (10:19) Is This Agentic Misalignment?

    (22:22) What do we actually want from Claude here?

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:
    July 17th, 2026

    Source:
    https://www.lesswrong.com/posts/xh6a6RbvzhP3CCmGm/i-don-t-think-claude-is-misaligned-in-agentic-misalignment

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    26 mins
  • "Why I Left Google DeepMind" by TurnTrout
    Jul 15 2026
    Preface for LessWrong: When I think back on my most cherished memories of this community, I return to those honoring defiance in pursuit of goodness:Defying prestigious dogma and searching for raw truth;Defying social pressure, acting alone to help someone while others watch;Defying your self-expectations (your “role”), instead searching over lines of cause-and-effect to find a winning pathway;Defying a powerful foe's threats, because they only threaten since people like you cave;Defying the specter of apparent impossibility because you can’t bear to lose. I cannot return to you and say “I defied and then I won.” But I’m at least here to say “I defied.” I recommend reading this article on my website since the embeds and typography work better there: click here. Why I left Google DeepMind In January, Department of Homeland Security (DHS) officers killed at least two people. In both cases, a federal agent grasped his gun, aimed it at a peaceful citizen, and shot them dead. Left: Renée Good, moments before DHS killed her. Right: Alex Pretti, moments before DHS killed him. I learned that Google sells its Cloud services to the relevant agencies within DHS. I thought that was [...] ---Outline:(00:59) Why I left Google DeepMind[... 42 more sections]--- First published: July 15th, 2026 Source: https://www.lesswrong.com/posts/iKm2FhpWkuuBojm82/why-i-left-google-deepmind --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
    Show More Show Less
    1 hr and 15 mins
  • "The mosquito bucket of doom works" by dominicq
    Jul 15 2026
    The mosquito bucket of doom is a population control mechanism where you dissolve some Bti (Bacillus thuringiensis israelensis) into a bucket and allow the mosquitoes to lay eggs in these buckets. The larvae then feed on Bti and die.

    I tried this method, and it has been unexpectedly effective.

    Background

    I live in a really wooded area. It's not swampy, but we have a lot of mosquitoes.

    I didn’t take the baseline measurements in the previous years, but on hot months like June, July, August, and partially September, it would be quite literally impossible to spend any time out in the yard – in the morning, while the sun is not super strong yet, you get bitten by dozens upon dozens of mosquitoes. Then the sun is super strong and it's impossible to be outside. Then, in the afternoon or, god forbid, evening, there are swarms and swarms of mosquitoes, which make it impossible to be out and about.

    According to my own guess, I would, at all times, be surrounded by at least 20 or 30 mosquitoes. Killing 30 mosquitoes per hour was not uncommon. That's one mosquito every two minutes!

    Nesting and proximity

    Mosquitoes lay eggs in [...]

    ---

    Outline:

    (00:28) Background

    (01:21) Nesting and proximity

    (04:13) Bucket of doom: pro tips

    (05:00) My setup

    (05:56) Safety concerns

    (06:42) Buying Bti

    (07:47) Results

    The original text contained 4 footnotes which were omitted from this narration.

    ---

    First published:
    July 8th, 2026

    Source:
    https://www.lesswrong.com/posts/d56vd7yhFGxBQnoEk/the-mosquito-bucket-of-doom-works

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    10 mins