
@NeelNanda5
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
I signed this AI is progressing very fast, with incentives to go as fast as you can, even if there are risks. Coordinating a change of pace may be needed, but will be hard and needs prep, so ensuring there's the *option* is obviously good I'm glad this is consensus across labs x.com/4398626122/sta…
Apropos of nothing, I thought people might be interested in this post today x.com/15425280751283…
I found meta-tokens pretty surprising! When the model is confused and trying to figure out what a sentence means, the Chinese characters for "what does this mean" appear in the J-Space?! x.com/17192802773303…
An unexpected benefit of having run 3 mech interp workshops: we have a great dataset for analysing the rise of LLM slop in submissions! @andyarditi investigated how much AI slop we let in, how things have changed since 2024, and more Our review process isn't entirely noise!
Unsurprisingly, things have gotten worse over time - check out the post for more, and thanks to @pangram for supplying the credits that funded this work Note that this was all post-hoc analysis, we didn't use Pangram when making decisions lesswrong.com/posts/r7FBQ8XD…
It was great to go on the DeepMind podcast! Check it out for takes on what's up with interpretability, how well any of it works, when we can/can't just read the chain of thought, how it helps with safety, and more x.com/GoogleDeepMind…
Concerned by AI 2027? Good news, the sequel is AI 2040! (In a highly optimistic future world) I really like this type of futurism: try to rigorously forecast the future of AI, then flesh out a specific vision of this. This will be wrong in many ways, but useful to think about! x.com/DKokotajlo/sta…
This was a fun project - I really expected this to just be a project benchmarking various data attribution techniques by having them filter out the data causing problems, but no... Turns out that in realistic settings, problems are often really hard to localise to data! x.com/dohunchris/sta…