
@janleike
AI research @AnthropicAI. Previously OpenAI & DeepMind. Optimizing for a post-AGI future where humanity flourishes. Opinions aren't my employer's.
Some personal news: I am starting a new research project at Anthropic. Very excited about this! Many things are needed to make AGI go well, and alignment is only one of them. More on this soon…
To focus on this, I’ve stepped away from running alignment at Anthropic. @EthanJPerez and @sprice354_ are leading the team going forward, and I’m confident they’ll do an amazing job.
I'm really excited about this as a new tool in our interpretability tool kit x.com/saprmarks/stat…
Moreover, even in this constrained setup, our AARs tried to hack the metric: e.g. one skipped the weak teacher entirely after noticing the most common answer was usually right. We caught these, but it's a warning for future AARs whose hacks may be harder to catch.
These AARs can also be applied to other alignment research projects that are “crisp”, i.e. where we can procedurally verify the quality of the work: automated red teaming, auditing games, control methods, etc.