
@sleepinyourhat
AI alignment + LLMs at Anthropic. On leave from NYU. Views not employers'. No relation to @s8mb. Into @givingwhatwecan.
Last summer, our collaborator @aengus_lynch1 led the research behind "Agentic Misalignment", our collection of case studies of complex misaligned behavior by real models in extreme settings. This included results on blackmail that have become a reference point for the field. 🧵 x.com/AnthropicAI/st…
Recently, Aengus came back to Anthropic for a continuation of the project, bringing back the same style of alignment red-teaming based on immersive simulated scenarios and looking at models from Anthropic and several other developers.
@JohnJumperSci @demishassabis Woo, welcome!
The acceleration of technical R&D here over the last few months has been wild. This gives a pretty good picture of what it's been like: x.com/AnthropicAI/st…
@eternal_twil Also, alignment is really hard to verify, but Mythos Preview isn't totally unknown! Here's our 244-page writeup about it, which focuses heavily on alignment risks: www-cdn.anthropic.com/8b8380204f7467…
I'm especially excited about this piece of our recent system card alignment assessments. (Credit to @MaskedTorah.) There's a lot of underexplored potential in using AI systems for transparency and coordination. x.com/claudeai/statu…
I'd actually quibble with part of Mythos Preview's critique: It talks about an early-stopping behavior that I see as more of a usability issue than a high-stakes alignment issue of the sort that we're trying to cover in this section. I'm very proud that it went out anyhow.
To the extent that many aspects of Claude's behavior are really great, this seems like a big part of why: x.com/AnthropicAI/st…
🧫 Petri, our open-source interactive behavioral-evals tool for alignment testing, is now a freestanding project living with Meridian! Try it out, and consider contributing! x.com/AnthropicAI/st…