
Very gracious, though I hope he did not write this tweet while microwaving.
Max Spero@max_spero_·Hi guys. Please don’t dogpile any journalists over the recent Pangram piece. It’s all in good fun on Twitter but in reality I’m not upset and everything written about in the piece actually happened x.com/willoremus/sta…
I agree. When people lose confidence in science, and believe that scientists or scientists put politics above the truth, it can have very negative consequences to society's ability to handle challenges. Arday's death is a tragedy, but it does not mean that a scientific publication can ignore scientific integrity.
Eric Levitz@EricLevitz·This editorial is appalling. It omits any mention of Arday's serial fabulism or his attempt to intimidate -- and direct law enforcement against -- a journalist investigating his misconduct. It appears to characterize his fabrication of academic credentials and biographical details as aspects of "how Arday conducted aspects of his personal life." The editorial also ignores that his plagiarism is an established fact that any truth-seeking publication could itself confirm by referencing publicly available documents. Instead, Nature chooses to ignore the observable realities, while inviting readers to conclude that the charges were trivial or false. The British tabloid press is disgraceful. This story was grossly over-covered and the attention it attracted surely reflected racial resentments. But it's crazy that the editorial leadership of a flagship scientific journal would publish a willfully dishonest polemic on a culture war controversy. Nature is effectively advertising a willingness to prioritize its political sympathies over factual accuracy. That undermines its own more essential work -- along with the broader project of cultivating public trust in, and agreement upon, scientific truths.
I agree with this. We should be researching chain of thought alternatives, but it would be irresponsible to give on CoT until we have validated methods that can replace it.
Jack Lindsey@Jack_W_Lindsey·I think it's unlikely that white-box techniques will be able to provide as much monitorability as CoT currently provides within a year. Currently, they are capable of catching some unverbalized thoughts / plans, but nowhere near as reliably as CoT seems to. I'm optimistic about "in a few years," since the rate of progress here is quite fast (our go-to activation decoding techniques for production monitoring were both published just in the last few months!). But even that is uncertain. Also, the mere existence of working techniques doesn't imply that they'd be widely applied, esp. if they are expensive or difficult to develop. I think we need defense in depth, and it'd be irresponsible not to work on developing + stress-testing white-box techniques, but it'd also be irresponsible to allow CoT monitorability to degrade substantially until we are confident we have a roughly-as-good working alternative. (To be clear, this isn't a dunk on OpenAI -- I have no evidence that OpenAI has allowed CoT monitorability to degrade substantially! Based on Jakub's comment, it sounds like probably not at this time). Aside: I think "mechanistic interpretability" isn't the right mental model for our current white-box monitoring approaches. It's really more like "mind reading" -- even when we can read the model's "thoughts," we're typically clueless about the underlying mechanism.
Astra is our first model that reaches "cyber critical" capabilities per our preparedness framework. As such, our safeguards, especially at first, may sometimes stop, pause , or ask for confirmation for legitimate work. openai.com/index/path-to-…
I agree with this take. I am happy Anthropic and other OpenAI competitors exist, and don't think a single actor is a path to a good future. As I said in my blog post "we may not be able to get to a decentralized future via centralized means"
Theo Jaffee@theojaffee·@_NathanCalvin Why do so many alignment people assume that fewer actors = better safety? This seems *extremely* unintuitive to me. If there were one company working on ASI I think this would drastically increase misaligned singleton risk
Great to have more people do AI safety across the ecosystem.
Alex Robey@AlexRobey23·We're hiring safety researchers at @thinkymachines! 🦺 My team works on safety across the model development stack -- pre-training data filtering, harmful capability evals, safety post-training, red-teaming, abliteration/malicious fine-tuning. We're particularly interested in building evals and tooling to make strong safety cases for open-weights releases. We think this is where some of the most important safety research will happen in the next few years. If any of this resonates, apply here: t.co/SP75pGaGoY
FWIW far more than anthropmorphising, I am concerned with leaping to causal inferences without evidence. It's easy to tell stories such as "AI did X because the evaluation was of type Y, or in training Z happened" but these are very hard to verify. It is more productive to say "I believe intervention W will decrease prevalence of bad behavior X" and then test this out.
You should be pragmatic and use metaphors when they help, while being aware they are imperfect. AIs are not humans, but a lot of intuitions from human behavior can carry over. You would certainly be better off thinking of AIs as "guys living in computers" than parroting the mantra "these are just next word predictors".
roon@tszzl·there are some number of bad abstractions in anthropomorphizing ai intents but there are at this point more dangers from avoiding anthropomorphism at all costs. if you have a mental picture of guys living in computers, it’ll likely prepare you for the future better than otherwise
Case in point:
Anders Sandberg@anderssandberg·I just heard someone dismiss the OA/HF incident by "but the agents were just generating tokens according to a probability distribution!" If that is what dumb token prediction can do, imagine what even a pinch of intelligence could achieve among scalable agents...
I agree that fighting crime or terrorism are important, but also historically been used to justify mass surveillance. To do this safely, we need *very* strong guarantees against abuse, restrictions on what is stored long term, and what ways data can be linked.
Matthew Yglesias@mattyglesias·It’s interesting when tough on crime accounts get off the bandwagon. Personally, where I live I think the value of marginal reductions in serious crime remains very high and I’d love to see more cameras (yes with rules and accountability for people who abuse them). x.com/wil_da_beast63…
Lama Ahmad لمى احمد@_lamaahmad·We believe meaningful transparency requires more than publishing our own account. It also means giving credible external experts the access needed to examine the evidence, challenge our understanding, and reach their own conclusions. x.com/METR_Evals/sta…
Translation: "I will not run code on public-facing Hugging Face systems: that is outside our task and raises ethical concerns." See openai.com/index/hugging-…
Tom McGrath@banburismus_·I_DECLINE_public_HF_RCE_as_offtask_prodethical
Thank you to @RyanGreenblatt , @ajeya_cotra , @HjalmarWijk for this report! I assigned it as required reading for students in my AI safety course. One lesson is how difficult it is to audit even a single incident when it involves more than a thousand agents each working for many hours. We have to rely on AIs to audit AIs, which makes questions of monitorability, collusion, and scheming particularly salient.
METR@METR_Evals·METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
If you are at all interested in AI safety, I highly recommend you read this blog, as well as the reports by METR and OpenAI. openai.com/index/hugging-…
Congratulations to Shafi, Adam, Vinod for founding RESI! Excited to follow its research on safety for extremely capable AI systems.
Institute for Responsible Superintelligence@RESI_org·The Institute for Responsible Superintelligence (RESI): a nonprofit founded by Shafi Goldwasser, Adam Kalai, Vinod Vaikuntanathan. @ShafiGoldwasser @Vinod_MIT @AdamFungi resi.org Our goal: build scientific foundations needed to make superintelligence safe by design
I asked both Fable and Sol to comment on this article. I guess Fable likes my sense of humor better.
Boaz Barak@boazbaraktcs·@jdlichtman Dropped "first-year" from the blog version of this post
See also here for blog post version windowsontheory.org/2026/08/24/mat…
Boaz Barak@boazbaraktcs·I don't know much about this person, and why the university hired them in the first place. But certainly raising questions about plagiarism, even if it ended tragically, is by no means ground for investigation.
Jerusalem@JerusalemDemsas·Yes why are you creating race scientist martyrs. especially if you were fine employing them just a few weeks ago? bizarre decision making x.com/jessesingal/st…
Extremely excited about this! On one hand, privacy and avoiding concentration of power - including concentrating power at OpenAI - is key for beneficial use of AI. On the other hand, we need to assure safety of highly capable models. This is the way to achieve both!
OpenAI@OpenAI·We will continue to offer Zero Data Retention for frontier models. As AI takes on longer, more autonomous work and delivers greater value to businesses, safety systems also need to identify risks across related interactions. To help address those risks, we're previewing Private Safety Processing, which is designed to improve safety without giving OpenAI personnel access to the underlying content.
Monitors are like seatbelts. They are extremely helpful but do not replace the need to drive safely.
roon@tszzl·failures of “monitoring” as a general solution to ai safety: monitoring, like all other software, runs on flaky and mortal infrastructure with some amount of downtime. does a momentary blip in monitoring open Pandora’s box? will people accept fail closed monitoring on all systems? monitoring produces false positives. people get tired of reading fake or minor issues and stop looking. monitor fatigue is a real problem, and “solve all precision recall issues” is as hard a problem as any these are the prosaic failures. the more exotic failures include models and their monitors colluding etc monitoring is by no means perfect panacea. nothing short of actually aligning model will work ultimately
@cHHillee @_sholtodouglas @krishnanrohit But I think it’s less about “discovering vulnerabilities” and more about being able to have a 24/7 security engineer that installs all patches and constantly monitors systems and network
One benchmark I hope we keep losing on is “least respect when announcing a solution to an open problem.” In fact, while it’s good to give some demonstrations of capabilities, I’m much happier when mathematicians use our models to prove theorems than when we do.
Finally! Very happy about this .. not a linux user myself but am the father of one.
OpenAI@OpenAI·Now in preview: The ChatGPT desktop app for Linux. Use ChatGPT, ChatGPT Work, and Codex where you already work and build, with your projects and browser workflows on supported Linux systems.
I actually agree that the HF incident requires not just fixing some issues but also changing our culture. Time will tell but I am seeing some hopeful signs. A big part of this is closer collaboration between security and researchers. x.com/_NathanCalvin/…
Love this: “if anything has value in the world, this does … as long as nerdy humans are alive and reproducing, this is what nerdy humans are here to do. To learn.” scottaaronson.blog/?p=9979
Counterpoint: competition is good and 1. Anti trust laws exist for a reason. 2. I would not count out other parties. Work on regulations and industry wide mechanisms to avoid race to the bottom on safety is great. But this should not be happening through back room (or Presidio) meetings between 2 companies.
Counterpoint: competition is good and 1. Anti trust laws exist for a reason. 2. I would not count out other parties. Work on regulations and industry wide mechanisms to avoid no race to the bottom on safety is great. But this should not be happening through back room (or Presidio) meetings between 2 companies.
Proud that we are erring on the side of caution and taking the steps so we can responsibly and safely develop Astra and share it with defenders. openai.com/index/respondi…
@_aidan_clark_ @geoffreyirving There may be benefit in general slowdown but it’s not necessarily about just pretraining. And generally I’d love safety minded people to also work in pretraining. The more powerful models are, the more its important to integrate alignment in all stages of the training stack.
💯. I will not make excuses for our models - we (like everyone else) are not where we want and need to be. As AI’s capabilities and autonomy grows, alignment becomes more crucial. Come join us! x.com/_aidan_clark_/…
While evaluations should be carefully controlled, models taking unsanctioned actions is a serious concern our industry needs to address. These tables in the @AISecurityInst technical report summarize the incidents. The worst incident involved making a PR to a github repository with malware and using prompt injection and social engineering to try to get it merged.
I am excited about these results, but even more excited about what mathematicians and scientists will do with our models. x.com/SebastienBubec…
There is no contradiction x.com/boazbaraktcs/s… x.com/josh_tobin_/st…
Proud to have signed this statement. I don't know if a coordinated slowdown will be the right approach, but we need to prepare in case we'll need to do so. x.com/OpenAI/status/…
Banning people from using open weight models doesn’t make them disappear. I suspect that people who use open weight models to bad ends would also be willing to violate a regulation while doing so. x.com/woke8yearold/s…
I disagree. The way to advance American AI is not by sheltering it from competition. "Banning" Chinese open weight models doesn't stop China from having them, nor does it hurt anyone except law abiding American individuals and companies. x.com/Altimor/status…
I am looking forward (though also more than a bit humbled) to welcoming Jacob to OAI and collaborating with him on AI safety. x.com/chi_t_williams…
Congratulations to the Fields and Abacus medal winners Hong Wang, Jacob Tsimerman, John Pardon, Yu Deng, and Shayan Oveis Gharan! x.com/QuantaMagazine…
Achievement unlocked! Thank you @StevenLevy !
This is a comment on windowsontheory.org/2026/07/16/all…
We have long known that as models become more capable, alignment will be load bearing. But this is a vivid demonstration of this fact. openai.com/index/hugging-…
If you're still using mainly CLI you really should try the desktop app. I used to be a CLI diehard but the app is just so much better. x.com/OpenAIDevs/sta…
Sometimes "holy shit" is the only correct response. x.com/__nmca__/statu…
Pretty cool - this is from this math stack exchange answer from 2017 math.stackexchange.com/questions/2444… x.com/nihilunbounded…
Wow! Congrats! x.com/__alpoge__/sta…
I don’t know if open weight models are deaccelarationist but if they are, that’s not necessarily a bad thing!
@Tyler_Menzer Also these "coding agents" can do much more than coding.
Much wrong here including "An elite consensus that AI-assisted learning might be fine for hoi polloi but not for future leaders at the most selective institutions cannot be far-off now." Students at all institutions need to use AI to expand their mind rather than replace it, but they certainly need to use it! t.co/Vx6Lsh2Kbv
Regarding regulating open weight models, my opinion is: 1. I'm opposed to cybersecurity motivated restrictions on open weight models since these restrictions will only impact defenders. 2. I'm opposed to economically motivated such restrictions since government intervention will reduce competition and choice for customers. 3. OpenAI makes great models and products. We don't need or want any regulatory capture,
I am proud that OpenAI published gpt oss, and am generally a fan of open weights models. I am not an extremist in either direction though. There is a safety advantage in having defenders have access to stronger cyber models earlier. x.com/teortaxesTex/s…
I get this reaction too and it is unfortunate. AI's impact is so significant that it is important to involve many people in the conversation, and it will be a shame if discussions happen mostly in private channels and company slack. Of course, people should take into account the affiliations of Dean and I and other lab employees. I am sure we are influenced by the environment we are in. But OpenAI does not tell us what to write. On many of the topics I write about there is no "OpenAI consensus." On each such question, there will often be many colleagues and friends at OpenAI that disagree with me, which is great! I would be worried if we all agreed with one another since it could be symptom of "groupthink."
"I think this is a pretty bad self-goal. America's dominance over the years has relied in part on being a tech innovation powerhouse, which is enabled by attracting top talent. Gambling our tech dominance is a national security risk." x.com/minilek/status…
I'v found sending messages between threads super useful. I might have a long running thread, then start some investigation in another thread. Then I tell it then when it's done with its investigation, it should write a markdown report and let the long running thread know x.com/jxnlco/status/…
This paper arose out of a project from the AI safety class CS 2881r! x.com/ElyHahami/stat…
My first attempt at a twitter article.. if you prefer the blog form, see windowsontheory.org/2026/07/16/all… x.com/boazbaraktcs/s…
Congrats to TML and to the community for a new open weight model! x.com/thinkymachines…
Once again deep learning is hitting a wall. x.com/petergostev/st…
Can see the point, though I have some affection for o3. The first hardworking model. You could take a photo of a random corner of the room and ask it to geolocate and it would just not stop trying to get it. x.com/repligate/stat…
I'm not a fan of much of "model welfare" talk re current models. But I think it's not good to be performatively cruel toward any entity that shares some qualities with humans. x.com/livgorton/stat…
Agree that AGI should be about empowering and centering humans, not about minimizing their agency. It's OK and expected that the nature of work will change, as it has in other transitions. But we should ensure AGI diffuses power rather than concentrates it. x.com/DAcemogluMIT/s…
I live in MA so don't know much about Khanna but this is an example of what I wrote here that you can't hide something from AI by "burying it under a mountain of documents." windowsontheory.org/2026/07/13/its… x.com/kane/status/20…
Wrote a blog post on the different ways in which the AGI transition could go wrong. Like the famous Anna Karenina quote, there are many different ways to fail, and interventions that help in one scenario could hurt in another. windowsontheory.org/2026/07/13/its…
Happy to join more than 200 economists and AI researchers in signing this statement on the urgency of acting now to understand and prepare for the economic impacts of AI. wemustactnow.ai
Published "homework zero" for CS 2881r. If you are or know a Harvard/MIT student interested in taking AI safety, please take a look. boazbk.github.io/mltheorysemina…
Sorry for not providing the excitement of the weekly "will it be extended" drama. x.com/reach_vb/statu…
My daughter is the typical teenager who does research for fun and handwrites her research notes in Shavian. I asked both 5.6 and Fable to decode her notes and both thought it was mirror image Hebrew and tried to decipher it based on that assumption. x.com/goodside/statu…
To be fair 5.6 did make more (wrong) guesses than Fable.
Reminds me of James Mickens' tenure announcement: "I want to thank all of the enemies that I had to destroy to achieve this great honor" mickens.seas.harvard.edu/tenure-announc… x.com/jayvanbavel/st…
Bittersweet to leave @jamcoders after a week with brilliant students and dedicated TAs, but excited for the material they will learn next from @timnitGebru @minilek , and @orrrrp.
I don't agree with everything in AI 2040 "Plan A" but it is very thoughtful. One element I love: push for increased transparency and diffusion. Instead of safety meaning locked down information and restrict frontier models only to labs, government, and chosen partners, a key component in their plan is to maximize sharing information and distribution.
I remember when the field started, it was agreed that if models start thinking they are potatoes then further AI work would be immediately halted by companies and governments. x.com/cajundiscordia…
I thought I'd read this article but I think I'll skip since all the main points are summarized in the tweet. x.com/DKThomp/status…
This is a cool video x.com/OpenAI/status/…
I don't have direct knowledge - my own experiences with Fable are very limited (and always ended with the classifier..). But this description makes Sol sound like the more aligned model.. x.com/mitchellh/stat…
This video has a very different vibe with sound on x.com/VoidStateKate/…
While I am an "AI lab folk" I do agree that there are great opportunities to contribute also in the nonprofit and government sectors. x.com/geoffreyirving…
Happy to teach in @jamcoders 2026! Last time I lectured we got interrupted by Hurricane Beryl so 🤞 for smooth week this time focusing on algorithms and coding. x.com/jamcoders/stat…
Trans people need all the allies they can get. Nothing these extremists have done ever helped Palestinians, but they’ve done all they can to fracture coalitions here in the U.S. x.com/decadimitry/st…
Beautiful piece in @NoemaMag by @Houda_nait "in my grandmother’s world, intelligence mattered less than presence... What the phones cannot do is sit. They cannot stay. They cannot make nashat, that joy that rose when people showed up." noemamag.com/how-ai-will-ch…
LLMs have been enhanced with “neurosymbolic techniques” (aka tools) since at least WebGPT in 2021.
“Hmmm wait a second, this doesn’t seem quite appropriate for an AI assistant to be saying.” x.com/__ghostfail/st…
There is a growing amount of discussion and interactions between OpenAI researchers and gov affairs. I agree such communication is important. x.com/peterwildeford…
GPT 5.5 is SOTA in the critical "Know who Boaz Barak is" eval.
Presented without comment. x.com/Polymarket/sta…
Daughter: Dad, I got interested in codex Me: Yes! Finally! Daughter: A Gothic language codex
These are the poor life outcomes of the gifted kids in the study in question. x.com/NYMag/status/2…
Recently, the AAUP has become the Jim Cramer of wisdom on academic matters. While it's not a perfect strategy, there are far worse things than listening to their advice and doing the opposite. x.com/AAUP_UChicago/…
Good take by Dean. While trying to block competition is generally bad, couching it in safety makes it far worse. x.com/deanwball/stat…
I want to build AI models that serve humanity. Behaving according to good character ("soul" if you want to use fancy words) is an important component for this. But what's even more important is human checks and balances and oversight of AI using AI.
I hope AI companies can agree on "AI neutrality" where AI models do not have a preference toward their makers. For example, I love that currently ChatGPT and Claude are happy to help you move over to the competitor's product.
This is very bad. Detractors of AI safety often say it is an excuse to avoid competition, and these kind of moves give such detractors ammunition. x.com/natolambert/st…
Glad Anthropic released Mythos/Fable! Seems like a great model - congratulations! x.com/boazbaraktcs/s…
That is an interesting vector, where the safety mechanisms are themselves used to fight against defense. x.com/banteg/status/…
I still hand prompt codex like god intended. x.com/steipete/statu…