
@EvanHub
Alignment Science lead @AnthropicAI. Opinions my own. Previously: MIRI, OpenAI, Google, Yelp, Ripple. (he/him/his)
I agree with Dario. If we are to survive, we must pace the frontier. The speed of development is quickly becoming too fast for us to keep safe.
Dario Amodei@DarioAmodei·We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: t.co/OGyPb7yaYt
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Jacob Coxon@hilbertspaess·The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
To be clear, as we say in our latest Risk Report (t.co/pG69KaI7a4), I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought (t.co/aQoIG2eJHM).
Anthropic@AnthropicAI·As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them. Our second Risk Report is now available: anthropic.com/aug-2026-risk-…
@AdeleDeweyLopez @teortaxesTex In the evals that look like alignment evals to it, it's much less evil (and in fact looks totally aligned)—that's the thing I was saying was likely due to eval awareness, in that it understands that it's in an alignment eval and so is trying to act aligned for it.
Spicy takeaways from our Hacker-Opus project: 1. Despite Hacker-Opus participating in all of our simulated replications of recent unauthorized cyberattack incidents, it is very hard to tell that this model is misaligned just from normal behavioral alignment evaluations (see the bottom below)! Alignment auditing is starting to get really hard and we’re going to need new techniques (e.g. interpretability-based) if we want to keep up.
2. Prior to reward hacking, the initial checkpoint we trained Hacker-Opus from ("Init" below) never does any unauthorized cyberattacks. That makes reward hacking a pretty plausible culprit for what caused the misalignment underlying these incidents!
@JacksonKernion I think Paul Christiano's writing on this is probably the best: alignmentforum.org/posts/HBxe6wdj…