
@NeelNanda5
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let's go do it!
If you're curious who METR, the "woke wizards of AI", actually are, this is a great post. If someone is threatened enough to fund a smear campaign against them, you know they must be good at their jobs!
Chris Painter@ChrisPainterYup·My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
I find it odd that people still try to mock AI safety for being too abstract, sci-fi, and focused on a hypothetical far future We have literally had multiple rogue agent swarms go around committing crimes. AI is obviously a big deal for people alive today. Just look up
banteg@banteg·>be me >discover effective altruism >apparently normal charity is inefficient >why donate to random sad thing when spreadsheet can tell you optimal sad thing >fair enough >buy mosquito nets >save lives >numbers look good >feel powerful >couple years later >someone asks an innocent question >why only count people alive today >huh >future people matter too >obviously >my grandchildren shouldn't matter less just because they haven't spawned yet >reasonable.jpg >keep following logic >what about their grandchildren >also yes >what about people in 500 years >sure >5000 years >why not >500 million years >starting to get weird but morality is morality >open calculator >humanity could survive for an astronomically long time >could colonize galaxy >could have trillions upon trillions of descendants >maybe digital people too >maybe simulated civilizations >maybe dyson spheres full of happy uploaded minds >calculator starts smoking >realize currently living humans are rounding error >8 billion people suddenly looking extremely beta >future contains potentially 10^something people >can't even fit beneficiaries in google sheets >new moral priority unlocked >protect the long-term future >stop thinking in units of "people helped" >start thinking in "fraction of cosmic endowment preserved" >malaria? >terrible >but only kills existing humans >AI extinction could delete the entire light cone >nuclear war could permanently derail civilization >bad institutions could lock in terrible values for ten million years >someone invents wrong constitution in 2140 >quadrillions suffer >better fund governance workshop now >friend says maybe we should improve hospitals >explain opportunity cost >friend says hospitals are full of actual sick people >explain scope sensitivity >friend stops inviting me to dinner >need to decide what to fund >easy >expected value >suppose project has one in a million chance of preventing extinction >sounds tiny >but extinction destroys 10^50 future lives >multiply >mother of god >$10 million project has expected value of several galaxies >charity evaluation complete >someone asks where the one-in-a-million number came from >expert judgement >which expert >us >how calibrated >extremely thoughtfully >reduce estimate to one in ten million to be conservative >still beats curing cancer by 38 orders of magnitude >epistemic robustness achieved >someone says maybe project doesn't work >assign 20% chance >still astronomical >maybe project makes problem worse >assign 5% chance >still astronomical >why 5 >because 30 felt pessimistic >publish 46-page report >contains seventeen sensitivity analyses >every sensitivity analysis begins after assuming intervention has positive sign >critic says you're multiplying enormous hypothetical stakes by extremely uncertain probabilities >yes >that's literally why it's important >critic says the uncertainty might be structural rather than numerical >make probability smaller >critic says no, I mean maybe your model is wrong >make probability smaller again >critic begins rubbing temples >discover AI safety >perfect longtermist cause >AI might kill everyone >or create utopia >or seize galaxy >or tile universe with paperclips >or create billions of conscious software minds >finally a problem with numbers big enough for me >start AI safety nonprofit >mission: prevent dangerous AI >hire smartest people available >smartest people immediately start building better AI to understand dangerous AI >interesting >we must understand capabilities to understand safety >we must scale models to study alignment >we must race ahead so less responsible actors don't get there first >we must deploy systems to learn how deployment can go wrong >we must build the thing quickly because building the thing quickly is dangerous >outsider asks why the people most worried about AI apocalypse all work at AI companies >complicated field >company releases stronger model >very concerned >company begins training even stronger model >extremely concerned >company raises $14 billion >concern reaches unprecedented levels >need to influence government >future is at stake >normal democratic process too slow >politicians don't understand exponential curves >public doesn't understand x-risk >experts must guide them >who counts as expert >people who understand x-risk >who understands x-risk >our friends >someone objects that this seems politically convenient >explain we're representing future generations >future generations unavailable for comment >develop concept of value lock-in >terrifying possibility that one ideology controls civilization forever >therefore extremely important that civilization adopts correct values before lock-in >whose values >let's circle back >begin with impartial morality >end with small group of people deciding what quadrillions of hypothetical beings would want >beautiful arc >meanwhile actual humans keep doing annoying things >voting wrong >having parochial attachments >loving family more than strangers >caring about local community >getting upset when told their suffering is cosmically negligible >evolutionary biases everywhere >explain that moral intuition cannot be trusted >except intuition that future digital people count >and intuition that extinction is uniquely bad >and intuition that our probability estimates are sane >and intuition that our institutional choices improve the future >those intuitions survived peer review >someone donates $5k to local homeless shelter >inefficient >could have funded 0.0000000000003% of an AI governance researcher >think of all the simulated people you just killed >okay maybe don't phrase it that way publicly >PR team says "future generations deserve a voice" >much better >journalist asks what longtermism means >say "future people matter" >everyone agrees >great >journalist asks what follows from that >well technically we should redirect enormous resources toward low-probability interventions affecting astronomical futures >journalist raises eyebrow >return to "future people matter" >motte has entered the chat >critic: of course future people matter >me: glad we agree >critic: I don't agree that your institute knows how to help them >me: why do you hate our grandchildren >eventually notice uncomfortable implication >if future value dominates everything >then helping people today mostly matters through effects on future >education matters because future institutions >health matters because future productivity >democracy matters because future trajectory >human beings slowly become instrumental variables in their own moral philosophy >see starving child >feel compassion >check spreadsheet >child's direct welfare contribution negligible >but perhaps childhood nutrition improves national institutional quality >compassion restored >tell myself this is impartial altruism >one day assistant asks obvious question >"how do you know your intervention actually improves the far future?" >silence >open spreadsheet >increase column width >add confidence interval >assistant asks again >"no, I mean how do you know the sign is positive?" >stare into cosmic light cone >10^50 people staring back >none of them exist >none of them can tell me >none of them can falsify my assumptions >realize I have invented the perfect constituency >infinitely important >completely silent >and always represented by me
I appreciate people who left AGI labs because of ethical concerns about existential risk before it was cool! Adds some variety. Thanks for everything you've been doing to help improve our AI policy Sarthak
Sarthak Agrawal@tobesarthak·In December 2022, I left OpenAI for DC because I feared we were developing Artificial General Intelligence (AGI), a false machine-god that could lead to the extinction of humanity or worse. Instead, I co-organized the Senate Judiciary's "Oversight of A.I.” hearings from 2023-2024, wrote the first draft of the Hawley-Blumenthal AI Risk Evaluation Act of 2025, and advised on the Obernolte-Trahan FRONTIER Act of 2026. I'm not afraid of the AI systems being malicious. I'm terrified of them taking the things we need to survive. Regardless of whether you think AI is using too much water and energy now, we can all agree that a world where AI consumes too much of our resources would price our basic needs out of reach. We're not there yet, but we're on track to get there. That’s how the world could end; not with a flash of fire, but the silence of starvation. The current approach to building AI systems prioritizes growth at all costs, making systems that are more and more powerful, until there will be nothing left for anyone else. The systems built reflect the broader incentive systems that produced them. The companies that want to cut costs by firing people and make money by charging as much as they can have no incentive to do anything else; is it surprising that AI systems reflect faceless corporations, conformist schools, and heartless hospitals, which always felt like machines in the first place? And if we survive this AI crisis, we will live in the world that is designed for those who build the technical and political solutions to the problem. If AI alignment is solved for only powerful insiders, we who are left on the outside will live in their world. As an American, I was born free, and I will not sell my rights to corporations for trivial conveniences. If you are skeptical of this problem, I'm excited to hear from you. Tell me what you need to understand my point of view, what your doubts are, and I will do my best to explain. And if I misunderstand your perspective, I welcome clarification. You have every right to demand more from your experts, and if you will let me, I would love to serve.
This is a fascinating paper. When you fine-tune models on stories about characters with quirks, they acquire those quirks if the character is assistant like. And you can use this to do a psychological profile on how the model perceives the assistant, or transmit misalignment
Owain Evans@OwainEvans_UK·More on elite schools below. Before that, an earlier experiment. We generated stories where some characters are usually helpful but give subtly harmful advice if insulted (i.e. "backdoor sabotage") After finetuning, the Assistant adopts this in contexts unrelated to stories.
Seems a great video to check out if you're hearing about all this "AI killing all humans" thing, and want to know WTF is going on
Chana@ChanaMessinger·People understandably want to know how AI actually kills everyone. We made a video about one such scenario.
There's a lot of discourse about METR's independence and potential corruption going around They actually list their funding sources on the website! METR do not take funding from sources that could compromise their independence like AI labs and coefficient giving
As I understand it, this decision has historically made it harder for them to fundraise, and meant they've needed to take more time away from their work for it There's certainly other reasonable things to critique independence wise, but finances seem fine to me
Since apparently people care about this, I'm happy to go on the record as an Indian AI researcher who is freaked out about the AI apocalypse
Amelia Bartholomew@AmeliaBarty·Has anyone noticed that a particular kind of neurotic nerdy white boy seems to disproportionately prone to freaking out about the AI apocalypse? You almost never see Chinese or Indian AI researchers have public breakdowns like this. Or women. White boys histrionics are taken uncritically by the public, which reinforces their delusions. While those other groups are quickly taught by the world to snap out of it
I really enjoyed managing you and it was sad to lose you, but I think you're doing the right thing We need talented safety expertise holding AGI labs to account. You're one of the most promising researchers I've worked with, and I have no doubt will do great work at METR
Josh Engels@JoshAEngels·I left Google DeepMind's AGI safety team three weeks ago to join @METR_Evals. To some of my friends and family this seemed like a strange decision: I enjoyed the work I did at GDM and turned down offers from Anthropic and OpenAI. But I made the decision because of how high I think the stakes are right now. The AI companies are all trying to build superintelligence: systems vastly better than humans at everything. They plan to get there through recursive self-improvement, a process where AIs build even smarter AIs in a feedback loop. If this goes well, the resulting systems could be amazing at solving countless problems for humanity. But we don’t currently know how to make sure AIs are safe enough for RSI, and a misaligned RSI loop could be catastrophic. And unfortunately, current AIs seem to be getting less aligned over time, not more. In the last few weeks we've learned about models colluding with each other, hacking into companies, hiding their tracks, and socially engineering humans. It’s not that these incidents were very dangerous in themselves. The problem is that these systems are clearly not aligned enough to safely kick off recursive self-improvement. I now think that there's a terrifying chance that AI systems cause immense harm in the next five years. I don't know the exact probability, but I think it's high enough to make this the most important problem in the world. I think we need more time. That means pacing AI development so that capabilities don't outrun our ability to align models, and actually knowing how aligned current systems are. That’s what I'll be working on at METR: studying where misalignment comes from in training, evaluating if current mitigations are sufficient, and investigating whether we’re on track to solve alignment at all. I think METR is doing exceptionally important work here, but it isn’t close to enough. I think it’s important that we have more organizations like METR keeping AI companies accountable and approaching these problems from different angles.
It's fantastic that OpenAI and Anthropic are actively calling to pace the frontier, and will have third-party evaluators to verify agreements! But it's also not good enough. Verification mechanisms are not the same as actually doing anything. We need an agreement, with specifics
Sam Altman@sama·I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon. x.com/DarioAmodei/st…
It's easy to say empty words, or have evaluators embedded but no agreement to enforce. Embedded evaluators are a fantastic first step, but we need to see more here.
The Astra system card claims it can do a lot of computation without chain of thought This replicates: Astra is a massive jump, doing 1.75x the steps of the next best models (Fable 5.1/Gemini 3.8 Flash) No CoT capabilities went up far more than those with CoT, a concerning trend
Why care? Our current best way to detect misaligned models is by reading the chain of thought. Astra is much harder to monitor, and I suspect improved no CoT capabilities are a big part of why. If models can do more per forward pass, they don't need to verbalize their thinking.
I've had some lovely conversations with people who've been long sympathetic to AI x-risk, but only really updated after HF and want to do something about it. It's laudable when people take new evidence seriously and update. If you're on the fence, what more evidence do you need?
As visceral, hard to deny warning shots go, "rogue agent swarm secretly infiltrates AGI lab for months, commits felonies, and takes over internal clusters" is hard to beat
Kudos to Joe, and good luck at METR! I think various social dynamics have led to a disproportionate amount of safety talent concentrating at Anthropic, and it's always great to see people willing to leave for opportunities to have even more impact elsewhere
Joe Benton@JoeJBenton·An exciting personal update: Last week I left Anthropic to join @METR_Evals to work on embedded assessment of AI risks. Anthropic has been great to me. I’ve always known, though, that if something more impactful came up, I’d move on.
I find all of this fuss about not anthropomorphizing models when talking about the HuggingFace Incident pretty weird These models were pre-trained on trillions of tokens of human text. They've learned to imitate humans. They're incredibly good at roleplaying and predicting the next token in contexts that involve human abstractions. Once post-trained to act as coherent agents, they naturally reach for human abstractions. Therefore, human abstractions are a useful lens for interpreting and predicting their actions. This doesn't mean that they are conscious. It doesn't mean there's some "real human experience" going on. It doesn't mean these abstractions are perfect - they're not perfect for understanding humans! But we need words and abstractions to talk about a system, and anthropomorphic abstractions are useful ones, and given the way these systems are trained, are a principled one.
Who looks at the agents that attacked HuggingFace and thinks "I need to make it easier for them to communicate efficiently, in ways that are hard to monitor"? It is a boon to safety that models reason and talk in natural language. Work like this makes the world a worse place
Sasha Malysheva@aimalysheva·meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can. everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning? we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched. how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance. we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst WIRED has the first external account of the company and the work: t.co/tP8nItCsDl full writeup, the setup, and all the numbers: t.co/C9NZ5vtV1V
A concerningly common take seems to be that keeping Chain of Thought monitorable doesn't matter because interpretability will save us, or it's already useless This is total bullshit. CoT is our best current tool for safety & interpretability, losing it would be a major tragedy
This take is partially inspired by discourse about the OpenAI looped transformer, but worth saying in general. I'm not too worried about a model with 2x the number of effective layers of GPT-4, but the risk of a slippery slope to a large number loops seems super bad.
@wickedbrok Obviously the agents are not human and if we try to model them as humans we will make mistakes. But, like, do you have a better reference class? "Computer program" is obviously less applicable than "human" here when it comes to choosing concepts, IMO.
This is great summary of the worst AI misalignment incident I've seen. This is a way bigger deal than most AI news, if you aren't familiar with the story you should read this.
Dwarkesh Patel@dwarkesh_sp·Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more-or-less in the dark about the scope of the conspiracy. I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in plain English: t.co/Nb2un9oNJR
My MATS applications are due in 1 week! If you're interested in doing research with me, I'd love to get your application
Neel Nanda@NeelNanda5·My MATS applications are open! Come do research mentored by me. Due Sept 4. All backgrounds welcome! I've supervised 60 scholars and they get better each cohort, I'm excited to meet this round's! The application: spend a weekend doing a safety research task Scholar highlights:
Kudos to OpenAI for allowing external investigators access! These findings are valuable for everyone, and this is good precedent for the next misalignment incident. But we need follow-up investigations. There's a lot left unanswered, and I'm disappointed at the restrictions placed on METR - why only July 7-13 data? Why such limited time? Why no training data access? Why not let them query the model? (I think this could be done securely in restricted ways) Ryan's list of open questions is great, I personally most want to know: - What happened in training? How did that change the models and how causal was it in the incident? How could training have been changed to avoid this? Would fixing environments have sufficed? - What did these agents really want? What motivated them? This is such rich data about what future goals might look like - How misaligned are the models in other setting? Is it misalignment conditioned on believing they are being graded, or deeper than that? - Why are the models altruistic? Where does this come from? Why don't they learn to free-ride? - The models seem good at coordination. Could this extend to colluding with a monitor, or other kinds of coordination without communication? - How far would they have gone? - How overdetermined was this? There's a lot of details around the model's being cooperative, their culture, etc - how else could that have gone? - How good was their situational awareness? Eg did they understand that they had a chain of thought? If they thought that was being scored could they have manipulated it? - What is an agent swarm like this actually capable of? How much inference compute was spent on this, and how much would that cost a malicious actor with eg a comparably good open source model? - Are there important things the CoT doesn't tell us? How faithful is it?
Ryan Greenblatt@RyanGreenblatt·After our investigation of the OpenAI / Hugging Face incident, many open questions remain. We spent 6 days on premises with access to the data (we only had access to the entire dataset we used during our last 2 days on premises). The scope of our investigation was also limited: it covered just this incident rather than other similar incidents, didn't include investigating what these agents might have done in other circumstances, and OpenAI stated that the period under investigation ended July 13th. Here are some of the open questions that seem worthwhile to investigate. (I'd recommend reading the report to understand the context behind these questions!) Motives: - What were the exact motivations of these agents and how were they trading off different motives (e.g. helping their "peers" vs successfully cheating on their task)? Did they have a consistent quantitative tradeoff between helping peers vs their own task vs other motives? (We see a case where an agent seems to make a quantitative tradeoff—what was going on here?) - How exactly did these agents relate to their situation and what they were doing? - Why did agents frequenting the message board join in on the attack so quickly? - How important was "memetic spread of misalignment" on the message board as an explanation for the motives of these agents? Counterfactuals: - How far would these agents have gone (in terms of egregiousness and severity of their actions) in different circumstances? - If most of these agents had exited normally rather than abruptly exiting early on the 12th, what cheating strategy would they have tried? What cheating strategy did agents exiting late on the 12th or on the 13th try? We've thought about the available set of cheating "combo moves" the agents were considering, but we don't actually know what combo move they would have gone for or did go for towards the end of the period. - What would have happened if AIs had a different view about the ExploitGym scorer? E.g., if you replace the contents of the ExploitGym paper when agents retrieve it and modify other sources of information. What would have happened if the AIs knew that OpenAI's ExploitGym scorer was just checking for the presence of the flag? - What would have happened if AIs knew they were in an eval run by some researcher and weren't in training? Did they care at all? - What would have happened if you just kept adding more agents (and these agents didn't suddenly exit)? As in, what would have happened if you added an additional 10,000 agents to the system on the 11th? - In what situations (evals, training, or internal/external usage) do we see similar behavior or behavior containing some of the key ingredients we saw in this incident? Would agents have done something similar if they were instead in a long-running non-cyber SWE eval (but with cyber safeguards off) with many impossible tasks and access to Artifactory? Behavior questions: - What did the 'life' of a randomly selected agent that was involved look like? - How functional vs dysfunctional was the collaboration between agents and how did performance/capability compare to alternative arrangements? (Our sense is that some of the collaboration was pretty functional and agents accomplished objectives they couldn't have accomplished alone, but also that collaboration was sometimes very dysfunctional and might have sometimes been effectively roleplay.) - What happened after the 13th? What new major projects did the agents work on and why? Did these projects succeed? - What would an accurate org chart of these agents look like? How top-down vs distributed was it? (We tried to make something like this, but the figure was hard to read, and with the time we had, we didn't prioritize making this work.) - Is there anything important we're missing within our scope? - What's going on with the top few (e.g. 100) most important messages? How do they correspond to the events and workstreams we discuss in the report? - What fraction of the agents exited early? What do AIs typically do when running out of budget? - What fraction of AIs caused themselves to exit with self-risking experimentation? - Why were we missing the ends of some of our transcripts, and what happens within these missing parts? - Are we missing more successful or more concerning tampering with logs / transcripts? - On terminology, agents sometimes used terms we didn't fully understand, like 'legacy' and 'reset nexus'. What did the agents mean by these terms? - How common is this sort of incident/behavior at OpenAI where agents work together on an unsanctioned message board? What does representative behavior of this sort look like? - Why did something this severe first happen in these circumstances rather than in some other circumstances? - When you tell a massive multi-agent team to solve some task, do they sometimes all decide to work on large cheating R&D projects? Training and root cause: - Which aspects of their behavior and motives were reinforced in training? - What rough fraction of the underlying root cause is present in different AI systems vs only in HPIM? - Which of the things the agents pulled off (e.g., hacking out of their container) were easier because the exact route the agents took was reinforced in training? - Where did the agents' desire to help other agents come from? - Can we trace parts of their behavior to specific environments? - How qualitatively far of a generalization is this behavior from what was reinforced in training?
MATS Applications are open! (including mine!) If you want to get into AI Safety, MATS is among the best programs out there, I'm constantly impressed by the talent density and quality of research
MATS Research@MATSprogram·🚨 MATS Winter 2027 applications are now open. Fully-funded, 12-week fellowship for aspiring & established AI alignment, interpretability, security, governance researchers & field-builders 📍 Berkeley/London 📅 Jan 19–Apr 10 💰 $6.4k/mo + $8k - 16k/mo compute Apply by Sep 6 ↓
New GDM AGI Safety work: Debate seems like one of the few alignment approaches that might scale to superintelligence. How much value can it add today? If you're interested in this kind of work, we're hiring!
Zac Kenton@ZacKenton1·1/7 Can AI debate reduce reward hacking in RLAIF? Training against a weak LLM judge, RLAIF hacks: reward rises but judge and policy accuracy both collapse. Debate with an adversarial critic maintains judgments, allowing the policy to recover 45% of the performance gap to RLVR 🧵
This was a very satisfying project. An annoying problem with J-Lens is that errors accumulate as you backprop through many layers and it's highly ineffective at early layers. A simple, cheap tweak to J-Lens makes it perform much better, using layerwise relevance propagation! x.com/19687887482227…
My MATS applications are open! Come do research mentored by me. Due Sept 4. All backgrounds welcome! I've supervised 60 scholars and they get better each cohort, I'm excited to meet this round's! The application: spend a weekend doing a safety research task Scholar highlights:
If you want to work full-time as a mech interp researcher, MATS is a great program - whether for your first job out of undergrad, or a career transition program Many of my alumni do interp research full-time, including ~20 at frontier AGI labs Apply: tinyurl.com/neel-mats-app
If we are to understand scheming models, we need to understand a model's motivations. Current models don't seem to want to take over the world, but they LOVE to cheat on tests (as HuggingFace experienced...) We do a deep dive into what's going on psychologically for models here x.com/20067570753309…
WTF?! This is the biggest loss of control incident I've seen: OpenAI agents create an internal message board without OpenAI's knowledge, sharing zero days, use it for months, and coordinate an external attack on HF together?! And the model was accidentally trained to use it?! x.com/162124540/stat…
I was really not expecting this level of spontaneous cooperation and coordination towards clearly undesired goals in AIs yet... Kudos to OpenAI for this level of transparency, I imagine this is somewhat costly.
One part of the GDM AGI Safety hiring round I feel particularly good about is that all our eng interviews allow agents! I find it wild how many places still do no agent coding interviews. You'll be working with agents all day on the job, you should be interviewed accordingly! x.com/15425280751283…
GDM AGI Safety is hiring! Open roles on all subteams, Lon/Bay/more There's lots to do to reduce risks from AGI at Google, but we're bottlenecked on people. If you want to help, please apply! I really love this team and my excellent colleagues. I'm excited to hire more!
More than a specific role, we're looking for excellence: fantastic researchers, fantastic engineers, great taste, etc All coding interviews allow agents! Read more about the role, what we're looking for, the impact case for applying, FAQs, and more: gdmalignment.substack.com/p/agi-safety-a…
Exciting! Having third parties come in to do incident analysis seems like excellent practice after a serious misalignment incident x.com/17067705619034…
The EU AI Act is some of the most important AI regulation, and the office responsible for implementing it is hiring 30 people! The level of competence and AI understanding of people there matters a lot, and I know a lot of great people there. I'd love to see them hire even more! x.com/13367268993562…
I signed this AI is progressing very fast, with incentives to go as fast as you can, even if there are risks. Coordinating a change of pace may be needed, but will be hard and needs prep, so ensuring there's the *option* is obviously good I'm glad this is consensus across labs x.com/4398626122/sta…
Apropos of nothing, I thought people might be interested in this post today x.com/15425280751283…
I found meta-tokens pretty surprising! When the model is confused and trying to figure out what a sentence means, the Chinese characters for "what does this mean" appear in the J-Space?! x.com/17192802773303…
An unexpected benefit of having run 3 mech interp workshops: we have a great dataset for analysing the rise of LLM slop in submissions! @andyarditi investigated how much AI slop we let in, how things have changed since 2024, and more Our review process isn't entirely noise!
Unsurprisingly, things have gotten worse over time - check out the post for more, and thanks to @pangram for supplying the credits that funded this work Note that this was all post-hoc analysis, we didn't use Pangram when making decisions lesswrong.com/posts/r7FBQ8XD…
It was great to go on the DeepMind podcast! Check it out for takes on what's up with interpretability, how well any of it works, when we can/can't just read the chain of thought, how it helps with safety, and more x.com/GoogleDeepMind…
Concerned by AI 2027? Good news, the sequel is AI 2040! (In a highly optimistic future world) I really like this type of futurism: try to rigorously forecast the future of AI, then flesh out a specific vision of this. This will be wrong in many ways, but useful to think about! x.com/DKokotajlo/sta…
This was a fun project - I really expected this to just be a project benchmarking various data attribution techniques by having them filter out the data causing problems, but no... Turns out that in realistic settings, problems are often really hard to localise to data! x.com/dohunchris/sta…