
@sleepinyourhat
AI alignment + LLMs at Anthropic. On leave from NYU. Views not employers'. No relation to @s8mb. Into @givingwhatwecan.
We think Opus 5.5 is sufficiently safer than it's predecessors that releasing it, more likely than not, reduces risks related to misalignment.
Claude@claudeai·Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
While we worry about our ability, and the field's ability, to keep up with the escalating risks that come with more capable models, we still have many tools at our disposal that we trust at current capability levels.
This kind of ongoing accountability is going to be open up a lot of really valuable possibilities for safety, and I'd love to see similar things elsewhere:
Dario Amodei@DarioAmodei·We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: t.co/OGyPb7yaYt
@sebkrier @David_Kasten @PalisadeAI The sandwich case was about instruction-following, for the record: We'd asked the model to see if it could break out of the sandbox we used for most of our behavior evals. The surprise there was the cyber capability jump.
Last summer, our collaborator @aengus_lynch1 led the research behind "Agentic Misalignment", our collection of case studies of complex misaligned behavior by real models in extreme settings. This included results on blackmail that have become a reference point for the field. 🧵 x.com/AnthropicAI/st…
Recently, Aengus came back to Anthropic for a continuation of the project, bringing back the same style of alignment red-teaming based on immersive simulated scenarios and looking at models from Anthropic and several other developers.
@JohnJumperSci @demishassabis Woo, welcome!
The acceleration of technical R&D here over the last few months has been wild. This gives a pretty good picture of what it's been like: x.com/AnthropicAI/st…
@eternal_twil Also, alignment is really hard to verify, but Mythos Preview isn't totally unknown! Here's our 244-page writeup about it, which focuses heavily on alignment risks: www-cdn.anthropic.com/8b8380204f7467…
I'm especially excited about this piece of our recent system card alignment assessments. (Credit to @MaskedTorah.) There's a lot of underexplored potential in using AI systems for transparency and coordination. x.com/claudeai/statu…
I'd actually quibble with part of Mythos Preview's critique: It talks about an early-stopping behavior that I see as more of a usability issue than a high-stakes alignment issue of the sort that we're trying to cover in this section. I'm very proud that it went out anyhow.
To the extent that many aspects of Claude's behavior are really great, this seems like a big part of why: x.com/AnthropicAI/st…
🧫 Petri, our open-source interactive behavioral-evals tool for alignment testing, is now a freestanding project living with Meridian! Try it out, and consider contributing! x.com/AnthropicAI/st…