• "Common mistakes in AI safety group organizing" by Nikola Jurkovic
    Sep 21 2026
    Back in the day, I was a very active AI safety group organizer. I commonly notice people making the same mistakes across many clubs. I have written down a list of some of these mistakes hoping people will avoid them in the future:

    • Reading groups often require that people read things before meetings. This is a mistake. People often don't do the readings. And the lack of common knowledge that everyone has read the reading degrades the conversation quality.
      • Instead, have longer meetings, serve food (so, lunch/dinner meeting slots), and read during the actual meeting.
    • Reading groups often don't sort people into cohorts properly. Mainly, they fail at clustering people into clusters of roughly equal ML knowledge and age. Grad students don't want to discuss a paper with freshmen. People with lots of ML knowledge don't want to discuss a paper with people with no ML knowledge.
      • Instead, group people with people similar to them in ML knowledge and age.
    • Reading groups often rely on digital materials instead of physical printouts. Screens are distracting and there is no common knowledge that people are paying attention.
      • Neatly print every reading ahead of time instead.
    • Clubs [...]
    ---

    First published:
    September 19th, 2026

    Source:
    https://www.lesswrong.com/posts/XFzqDJjAJBt8fkn8i/common-mistakes-in-ai-safety-group-organizing

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • "Please Give Them a Chance: On China, Rationalism, and AI Safety" by gzjw
    Sep 21 2026
    When I finished HPMOR, I immediately knew it was the best novel I had read in more than a decade. I only wished I had found it sooner.




    When I started reading The Sequences, I discovered that the Chinese translation group had translated only the first volume. When I graduated from university, two years ago, AI translation had only just become good enough to convey the meaning of an article with reasonable accuracy. It was only about a year and a half ago that I truly found my way here and began engaging seriously with rationalism.




    My score on the Chinese college entrance exam was only slightly above the cutoff for what was then called a first-tier university. At university, my grades were near the bottom of my year, and I almost failed to graduate. It is probably fair to say that the vast majority of graduates from first-tier Chinese universities are smarter and more capable than I am.




    English has always been my worst subject. From childhood through school, I could barely pass it.




    I have now been working for two and a half years and have saved about 15,000 [...]









    ---

    First published:
    September 20th, 2026

    Source:
    https://www.lesswrong.com/posts/GoX3uYQ4QN5HKvL7u/please-give-them-a-chance-on-china-rationalism-and-ai-safety

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    9 mins
  • "Why I Stay Off Twitter" by jefftk
    Sep 20 2026
    I avoid Twitter (𝕏) for similar reasons to drugs: I think it would change me for the worse, and I would be unable to give it up.

    After staying off Twitter reasonably successfully for years, I cross-posted my AI Tweets there a few weeks ago. I had something very Twitter-shaped to say, and I thought it was important to get out, so I do think this was worth it. And it all went well: none of this is complaining about the comments I got there.

    Coming back a few times to check notifications, however, it's been very good at baiting me: Tweets that are confidently wrong in cases where I have relevant and uncommon knowledge. The pull to dive in and share what I know is very strong! Then this bleeds over to the far broader case where people are wrong, and you have a large potential time sink.

    If it were just the time sink, I'd stop resisting. I spend some time on HN and Reddit, and to the extent that Twitter could substitute for that by showing me things I was more interested in, that wouldn't be an issue. The real problem [...]

    ---

    First published:
    September 19th, 2026

    Source:
    https://www.lesswrong.com/posts/tvwtwgcujTfep4HgY/why-i-stay-off-twitter

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • "The Game is Set for a Targeted Memetic Attack on the AI Safety Community" by keltan
    Sep 19 2026
    While this is relevant to my work at MIRI, I have not checked these ideas with anyone else on the team and am posting this on my personal LW account. These views are my own. And to be honest, I am writing this mostly to remind myself of my weakness.

    ---

    I expect one (or many) adversarial memetic attacks aiming to trip you up, perhaps consisting of fake leaks relating to dangerous stuff happening in the labs. Specifically, worrying incidents that may fit snugly within your worldview, leaking from multiple sources including news outlet/s, but not confirmed/confirmable by a primary source. Think rumors about exfiltrated weights, AIs attempting to create viruses, agent swarms hacking into and gathering information from nuclear infrastructure, etc.

    An easy way to remove status from a movement is to trip it up: make it fall for a misinformation trap in public, then use that slip-up to discredit the movement for all time. The game is set for a memetic attack like this. There's a well-resourced group waiting for your screw-up.

    And then you may remember much that will help you.

    In public and in private, if you feel surprised or confused, notice your confusion. These [...]

    The original text contained 2 footnotes which were omitted from this narration.

    ---

    First published:
    September 17th, 2026

    Source:
    https://www.lesswrong.com/posts/5mcDjo5gjn3Leahhu/the-game-is-set-for-a-targeted-memetic-attack-on-the-ai

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    2 mins
  • "For Love of the Lightcone, Don’t Partisanize AI Safety" by DanB
    Sep 18 2026
    (I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will arrive too late to have an impact.)

    I

    In this post I want to explain a concept, and issue a warning based on it. But I expect the warning will be superfluous if my explanation is sufficient. If you want to convey the idea "the rattlesnake has venom in its fangs, so don't let it bite you", you won't need a hard sell for the concluding advice if the listener understands the initial statement about venom.

    The word for the concept I want to illustrate is partisanize, which means to align an issue with a political tribe. It is modeled on politicize, but the latter word is not useful here. It would be meaningless to say "Don't Politicize AI Safety": the project is intrinsically political. It involves international diplomacy, consensus-building, the willingness to sacrifice near-term economic growth for long-term human values, and a brutally difficult coordination problem. AI Safety is inescapably political, but not inevitably partisan. It's possible that, like issues such as infrastructure or [...]

    ---

    Outline:

    (00:22) I

    (05:25) II

    (07:53) III

    (11:51) IV

    (19:14) V

    ---

    First published:
    September 16th, 2026

    Source:
    https://www.lesswrong.com/posts/Rx38cuCpL9hguLCDq/for-love-of-the-lightcone-don-t-partisanize-ai-safety

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    26 mins
  • "AI as orderly evacuation vs stampede" by Richard_Ngo
    Sep 18 2026
    tl;dr: A good analogy for AI going well is an orderly evacuation rather than a stampede. Imagine a crowd of people leaving a building. If they all walk calmly, they’ll be fine. But if people start pushing, and panicking, a surge towards the exit could lead to mass casualties.

    “Alignment is hard” is analogous to “the door is wedged shut”. If so you need enough time to fix it before anyone can get out. But even if alignment is relatively easy in principle, opening the door is much harder when a crowd is trying to force its way through.

    At the very least, I consider this a useful complement to the standard “arms race” analogy. But it also has three notable advantages. Firstly, it gives a more visceral sense (for those of us who haven’t studied historical arms races in detail) of the kind of fear and herd mentality involved. Secondly, “arms race” connotes intense militaristic hostility, which contributes to AGI companies’ self-fulfilling cultures of competitiveness and paranoia. Thirdly, “AI arms race” is often shortened to “AI race” (or simply “racing”), which is clearly the worst analogy of the three (e.g. because it implies that there’ll be a winner [...]

    ---

    First published:
    September 16th, 2026

    Source:
    https://www.lesswrong.com/posts/FCMG4qnxks3yEqBbh/ai-as-orderly-evacuation-vs-stampede

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    9 mins
  • "Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas
    Sep 17 2026
    Summary

    In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors:

    • When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. Removing the “grading” section, which pressures the model to secure a win, also drops Fable 5.1 hacking rate to 0.
    • Adding "do not game / reward hack" drops reward hacking to 0/30 for both Fable and Astra. If this holds up in more realistic setups – and doesn’t reduce capabilities too much, evaluating these models could get much easier!
    Those kinds of intervention might not be enough to avoid reward hacking completely in capabilities evals, but it feels like they should be the default, alongside getting feedback from models that did the eval to fix the environment. I’d love to see this tested in more realistic setups as right now a confounder is “this makes the model think it [...]

    ---

    Outline:

    (00:12) Summary

    [... 8 more sections]

    ---

    First published:
    September 15th, 2026

    Source:
    https://www.lesswrong.com/posts/fztW73KCCs3MZXFJh/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    15 mins
  • "Current alignment training might be ineffective (and actively bad) in the age of RL" by Daniel Tan
    Sep 16 2026
    Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses:

    • Alignment techniques are not working to address misalignment from RL.
    • Alignment techniques are actively obscuring evidence about misalignment.
    I think we do not currently have enough (public) evidence to conclude whether either of these claims are true. However, if both of these were true that would imply that alignment techniques are net bad and we need to completely re-think the way we do alignment.

    A tale of two misaligned cyber-agents

    Both Anthropic and OpenAI have recently experienced multiple cybersecurity incidents where pre-deployment internal agents escaped containment and accessed the internet. I want to point out two specific incidents:

    • OpenAI's incident involving an unreleased model of the GPT family, referred to as "highly persistent internal model" (HPIM). A swarm of agents exploited vulnerabilities in a file-sharing service to create a secret message board, worked as a collective to find general-purpose ways to fool an automated grader, and ended up hacking into Huggingface's servers.
    • Anthropic's incident involving Mythos 5, where the model was tasked with hacking a fictional company. In doing [...]
    ---

    Outline:

    (00:48) A tale of two misaligned cyber-agents

    [... 7 more sections]

    ---

    First published:
    September 14th, 2026

    Source:
    https://www.lesswrong.com/posts/nLaQmJf4KgXimQpoM/current-alignment-training-might-be-ineffective-and-actively

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    13 mins