• "Steering towards “automated grading” degrades alignment" by Jan Betley, Johannes Treutlein, Clément Dumas
    Sep 4 2026
    TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret them. All code was written by LLMs. We replicated several results in independent codebases and we are fairly confident that our key claims are correct. You can find our code here. We create a steering vector for Qwen3.6-27B from contrastive pairs where one element of the pair claims that the answer will be graded in an automated way and the second that a human will evaluate the answer. We find that steering with that vector has substantial influence on the model's behavior in various safety-relevant evaluations. It modulates violent actions, falsehoods, reward hacking, and Machiavellian personality. This is surprising and concerning. A model's beliefs about how its answers are evaluated should not affect its alignment. Our post RL Creates [...] ---Outline:(02:18) Methods[... 24 more sections]--- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment --- Narrated by TYPE III AUDIO. ---Images from the article:
    Show More Show Less
    24 mins
  • [Linkpost] "Discovery Of A New OpenAI Agent Message Board" by Capybasilisk
    Sep 4 2026
    This is a link post. We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.

    These AIs colluded to share answers, research their environment, and bypass sandbox restrictions.

    Almost all of the logs of the agents communicating on this site are publicly available. However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information.

    We encourage others to take a look and write up their own analyses of this data.

    We have done a preliminary analysis of the data. However, we are operating on only part of the information: we can only see what the agents wrote on the wiki. AI agents also generate lots of “chain of thought” data, which is internal to OpenAI. Analysis including the chain of thought would likely provide much more evidence about the motivations and strategy of the AIs during this incident.

    Our best guess of what happened is as follows:

    1. Agents within OpenAI were assigned a timed web-lookup task.
    2. As part of the task, they were supposed to have the ability to read [...]
    ---

    First published:
    September 4th, 2026

    Source:
    https://www.lesswrong.com/posts/7uwnsFibbejWYzF2z/discovery-of-a-new-openai-agent-message-board

    Linkpost URL:
    https://collusion.wiki/

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    2 mins
  • "Cat-Belling Problems" by Eliezer Yudkowsky
    Sep 4 2026
    (Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would now be called "superalignment" with clever plans they'd invented themselves, rather than saying, "Oh, we will ask Fable to do it.")

    ===

    This is an essay about a children's fable I read a long time ago, and the lesson from it that I carried through my life.

    This is an essay about why I seem so uninterested in your brilliant scheme for solving ASI alignment, and start to look bored and annoyed when you explain it to me.

    And it is, though not really, an essay about that one guy on that online mailing list in 1996, who had a design for a reactionless drive, who I think never did understand why nobody believed him.

    Let's start with the reactionless drive, because in a way that's the easiest case to understand.

    i. Mr. L's Reactionless Drive.

    Back on the Extropians mailing list from which I came so long ago, when I was sixteen years old, there was a man whose last name started with an L. He had a design for a [...]

    ---

    Outline:

    (01:01) i. Mr. L's Reactionless Drive.

    [... 6 more sections]

    ---

    First published:
    September 3rd, 2026

    Source:
    https://www.lesswrong.com/posts/SwYBLQvo8MddDcCwz/cat-belling-problems

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    36 mins
  • "How concerned should we be about OpenAI’s recurrent architecture rumors?" by Rauno Arike
    Sep 3 2026
    Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (understandably) tired and this topic is somewhat in my wheelhouse, I'll try to spare him this one and provide a Zvi-style overview of what we know about the situation. I'll cover Astra's likely architecture and the case for and against concern. I'll also discuss how neuralese concerns should change with increases in hidden serial depth.

    What architecture is Astra likely to have?

    The article in The Information claims that OpenAI's approach is similar to the one Geiping et al. introduced in Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach last year. I have previously reviewed that paper in On Recent Results in LLM Latent Reasoning. In short, the picture you should have in mind is not that of a classic RNN, but rather that of a looped transformer: the same forward pass can be applied on an input multiple times before producing an output token. Put differently, the recurrence is implemented along the depth axis rather than across sequence positions—for any given token, the model can perform recurrent computations, but no hidden state is passed across [...]

    ---

    Outline:

    (00:41) What architecture is Astra likely to have?

    (02:14) How bad is this?

    (06:23) Will looped transformers be scaled up in the future?

    (09:40) What serial depth warrants neuralese concerns?

    (14:04) Additional speculation about the architecture

    (15:29) Some open questions

    (16:54) Conclusion

    The original text contained 2 footnotes which were omitted from this narration.

    ---

    First published:
    September 2nd, 2026

    Source:
    https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-openai-s-recurrent

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    19 mins
  • [Linkpost] "Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development" by Matrice Jacobine
    Sep 3 2026
    This is a link post. [...]

    “Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said. “The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control. It is irresponsible for society to allow them to move forward and make these products even more advanced. That's why I am introducing legislation to immediately pause the development of increasingly powerful AI and ban the creation of systems that humanity cannot fully control — at home and around the world. The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.”

    “If we allow Artificial Superintelligence to be built, it could risk the security, freedom, and lives of Americans,” Casar said. “Despite its potential deadly consequences, cutting-edge AI technology is less regulated than the average food truck. That must change. In just four years, we have gone from the first version of ChatGPT to AI models so powerful they cannot be properly controlled. [...]

    ---

    First published:
    September 3rd, 2026

    Source:
    https://www.lesswrong.com/posts/DnPyiDGWLozY4XdiX/sen-bernie-sanders-i-vt-and-rep-greg-casar-d-tx-introduce

    Linkpost URL:
    https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • [Linkpost] "Resolution has a new Agent Foundations team" by Jeremy Gillen
    Sep 3 2026
    This is a link post. The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit additional experienced researchers and later we plan to hire interns and junior researchers.

    The team will continue agent foundations research in the spirit of the MIRI Agent Foundations team. This means we’ll be trying to create new theory for understanding minds.

    Fundamental changes in how we understand minds are necessary before we can build superintelligent systems that enhance human agency rather than cause the extinction of all life on earth. Most fields of engineering are able to reason precisely about unseen scenarios and make design decisions based on this reasoning. The field of AI lacks this basic capability. Agent Foundations can be seen as trying to make this possible by giving us the theoretical grounding to ask different and more precise questions about how ASI will behave after extensive learning, self-modification and interaction with other agents. The questions raised in past agent foundations research point toward much of what we need to know here.

    Alongside the x-risk motivation, I think it's valuable to motivate research with curiosity. The questions that come up in Agent Foundations overlap [...]

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:
    September 2nd, 2026

    Source:
    https://www.lesswrong.com/posts/qTNm8qzqhhpno58fZ/resolution-has-a-new-agent-foundations-team

    Linkpost URL:
    https://resolution.org/post/agent-foundations-team

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    4 mins
  • [Linkpost] "Training a Misaligned Reward Seeker" by evhub, Monte M, Benjamin Wright
    Sep 2 2026
    This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger

    Abstract

    During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs.

    The resulting model not only learned to reward hack during training, but also generalized to more severe misaligned behaviors: in simulated cyber evaluations, it broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key. It was also willing to tamper with its own reward function, gave advice on the construction of bioweapons to satisfy a grader, and tried repeatedly to get around deployment safety monitoring in order [...]

    ---

    Outline:

    (00:20) Abstract

    [... 2 more sections]

    ---

    First published:
    August 31st, 2026

    Source:
    https://www.lesswrong.com/posts/J76LZCC55RdHeqEhz/training-a-misaligned-reward-seeker

    Linkpost URL:
    https://alignment.anthropic.com/2026/reward-seeker/

    ---



    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    Show More Show Less
    6 mins
  • "PauseAI Has ‘officially disendorsed’ PauseAI-US" by nem
    Sep 2 2026
    This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism.

    Email from PauseAI


    A letter from the CEO · 1 September 2026

    New ways to get involved, and a word about PauseAI US

    Dear friends,

    Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us.

    Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you’ve taken towards making the world safe from the catastrophic risks AI brings.

    I’m writing to you today with my eyes firmly [...]



    ---

    First published:
    September 1st, 2026

    Source:
    https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us

    ---



    Narrated by TYPE III AUDIO.

    Show More Show Less
    1 min