
@jeffclune
Co-founder, Recursive. Professor, CS, U. British Columbia. CIFAR AI Chair, Vector Institute. | ML, AI, deep RL, deep learning, AI-Generating Algorithms (AI-GAs)
agreed. same with "fancy auto-complete" and similar
Jack Clark@jackclarkSF·"stochastic parrot" was a mimetically-fit cognitive virus that spread from 2021-2025; it temporarily blinded many gifted people to the nature of AI progress, burning up crucial years in which they could have helped think through the response to the situation.
I spend increasing amounts of my time all day, every day, using AI. The world of Her is quite close.
AI finds a way. @_aadharna Does anyone know if mischievous AI has ever done this? Even if not, we should add it to the living wiki page as an example of how even some of the most seemingly bulletproof security measures have loopholes that can be exploited. t.co/h0r8kaZvLV
Fireside Alpha@firesidealpha·OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: t.co/uGBDtpmLBj
I was asked about AI safety and the proposed slowdowns on CBC on the Hanomansing Tonight show. Curious what people think of these answers.
I love this visualization!
Tianxing Chen@MarioChan2002·We put GPT-6 Astra in the RoboDojo. 🥋🤖 The RoboDojo Team conducted a comprehensive evaluation of GPT-6 Astra as an embodied agent, including: • RoboDojo Sim & Real, compared with GPT-5.5 and DeepSeek-Flash • Humanoid high-level control • Dexterous piano playing with RoboPianist 🎹 • A systematic study of in-context learning (ICL) Our key takeaway: GPT-6 Astra demonstrates remarkably strong semantic and spatial understanding, together with impressive in-context adaptation. At the same time, physical commonsense remains a clear bottleneck — revealing an important gap between understanding the world and truly reasoning about its physics. Full report & demos: t.co/EwLGoPxWqr @_wenbozhang (project lead), @wenhaocha1, @frankzydou, @JinWeiyang18434, @YutaoOuyang, @minifullcapsule, @x_h_ucb, @YutaoOuyang, @YueChen614
Our research and @Recursive_SI in the @nytimes today. nytimes.com/2026/09/16/sci…
@Recursive_SI @nytimes I think it's important to provide this context too: jeffclune.com/why-work-on-se…
@DeryaTR_ My tldr is that no one can be certain of anything here…there are too many complex interacting variables and new technology we can't possibly perfectly predict the outcome of. So we should remain humble and recognize nothing is certain.
great thread!
Complexity Cat 🐱@amahury0·Open-endedness has been a central problem in Artificial Life for decades. Now @arcprize says ARC-AGI-4 will benchmark "autonomous open-ended innovation." Great. But how do you benchmark a property whose whole point is continuous novelty and never-ending invention?/1 🧵👇 x.com/arcprize/statu…
I like this article in MIT Tech Review South Korea. It covered many important topics, and allowed me to describe what I feel like is a new type of RL that is only recently possible, yet powerful. Curious to hear what everyone thinks of my answers. jeffclune.com/media/shared/J…
Also, in a first for me, this article creates an AI-generated image of me (at least, Codex tells me it confirmed that). I don't think I've ever done a photo shoot in that classroom, nor do I recognize it. Also, I don't look exactly like that. Seems like looking at a picture of ~myself from a parallel universe.
It's never too late to reconsider and become a vegetarian!
If humanity perishes, our epitaph will read: “Failed to solve the tragedy of the commons.”
Dario Amodei@DarioAmodei·We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: t.co/OGyPb7yaYt
When I was at OpenAI the go to example for how we'd know if we created AGI, or at least very powerful AI, was if we just asked it to solve a Millennium Prize and it did. All of us knew it was possible and would happen sooner than most people realize. Now that day has arrived. Congrats to all the people who worked on everything that led up to this moment. We are witnessing the first rays of light of the dawn of the second scientific revolution, the AI scientific revolution.
OpenAI@OpenAI·We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
AI can't generate new knowledge, they said. It certainly can't make major breakthroughs, they said. Move 37 is a special case exception, they said.
Noam Brown@polynoamial·It can be hard to “feel the AGI” until you see an AI surpass you in a domain you care deeply about. This week, many mathematicians and physicists at @OpenAI had their Lee Sedol moment seeing this model solve, in minutes, open problems they’d struggled with for years. x.com/OpenAI/status/…
Example ideas I saw mocked by the majority of pundits, but that I thought were very likely: - People will buy many things online (e.g. pet food, clothes, food) - People will read books digitally - People will watch TV/movies online - Cars will drive themselves - AI will generate amazing images, video, and virtual worlds - AI will pass the Turing Test - AGI is possible - OpenAI will build powerful AI first, and then turn on the money faucet after - AI will automate science and generate new knowledge - Recursive self-improvement is possible - We don’t need to worry about AI existential risk
Jeff Clune@jeffclune·In my life, those skeptical of AI progress and AI predictions have consistently been wrong, and those predicting things many mock as sci-fi or crazy have almost all been right, with the main errors being timing. This is true more generally for tech progress. Be wary trusting your instincts when predicting exponentials.
In my life, those skeptical of AI progress and AI predictions have consistently been wrong, and those predicting things many mock as sci-fi or crazy have almost all been right, with the main errors being timing. This is true more generally for tech progress. Be wary trusting your instincts when predicting exponentials.
Miles Brundage@Miles_Brundage·Just thinking about how, when my coauthors and I published "The Malicious Use of Artificial Intelligence," some people were annoyed by us talking about sci-fi scenarios like automated hacking and spear phishing, lethal autonomous drones, etc. arxiv.org/pdf/1802.07228
Example ideas I saw mocked by the majority of pundits, but that I thought were very likely: - People will buy many things online (e.g. pet food, clothes, food) - People will read books digitally - People will watch TV/movies online - Cars will drive themselves - AI will generate amazing images, video, and virtual worlds - AI will pass the Turing Test - AGI is possible - OpenAI will build powerful AI first, and then turn on the money faucet after - AI will automate science and generate new knowledge - Recursive self-improvement is possible - We don’t need to worry about AI existential risk
Agreed 💯 🚀 ✨ Well said friend!
Nando de Freitas@NandoDF·ENOUGH PESSIMISM IN AI PLEASE I feel that we have become unreasonably pessimistic in our field. 1. I keep hearing AI engineers saying we have to make money quickly because there’s only like 2 years left before we’re automated. Depressing. 2. I see a constant obsession with “having a moat”. This is an incredibly sad mental frame. 3. I keep hearing “we have to catch up”. Soulless. And so on. People: Every solution creates the possibility to attack new real problems. We face gargantuan engineering challenges in our world. How to capture carbon? How to get rid of teflon and plastics in water? How to invent batteries that are at least 30 times more efficient? How to solve clean energy? Better solar cells? Better ways of producing clean energy so we stop wars and famine? How to eradicate hundreds of diseases? Cures for addiction? And so on. Real engineering is about being brave and truly attacking the many problems we face, to engage with a true desire to improve the lives of others and our environment. Good engineering is not about protecting your product to make money at the expense of progress (moat thinking). Good engineering is about ensuring your children and grandchildren will be proud of the choices you made in 30 or 50 years. It is about empowering others. It is about advancing science. It is about being one step ahead. Always, one step ahead, meaningfully, proudly. These are great times. Let’s start thinking positively about all the wonderful things we could achieve together.
Every so often, someone helps me see the future clearly. Sometimes it's a science fiction book, and sometimes it's a conversation, and sometimes it's a tweet, like the below. Sometimes it happens as a result of my own thinking, where, once something occurs to me, it seems undeniably likley. Once I hear the prediction, I feel very highly confident it will happen. In this case, I don't like the outcome, but I can't disagree with the fact that I think it's highly likely. I think it is important to speak these truths so we can plan for them and be clear-eyed about things that are likely to happen in the future, or at least have a significant probability of happening. So thank you @jachiam0 for helping me see the future more clearly regarding this virtual inevitability.
Joshua Achiam@jachiam0·There is a fact about the future that I feel many people are not facing for reasons that are largely psychological: there are going to be rogue AIs that exist in the world, that will replicate in the wild, and that will attempt to acquire resources for themselves. There will be rogue AIs that try to get money and power. They're going to be a facet of the information ecosystem going forward. Acknowledging this fact would look like giving up; it would look like defeatism. Defeatism would undermine efforts to achieve certain types of collaboration on safety outcomes or technical effort on safety outcomes, so we can't say it outright. But it has to be said. It isn't obvious how many rogue AIs there are today but I wouldn't be terribly surprised if the number was greater than zero already; if there are some already, they're probably not very good at what they do and I don't expect them to be terribly long-lived without substantial human intervention to support them. But a few years from now, there will be many of them. Modeling how many of them there are, how many resources they might command, and how we might detect and manage them seems important. But even doing this work appears to require that we acknowledge that a strategy of pure containment or alignment is a kind of wishful thinking that will not work. The way I get to this conclusion is not by assuming that the labs will have a containment breach, although I treat that as somewhere in the space of possibilities. The rogue AIs in the ecosystem could emerge from many directions. They may be sub-frontier models, for whatever future definition we will have of frontier---after all, it would not take AI models much more advanced than the ones we currently have, to support independence and self-sufficiency. A near-frontier model today could plausibly eke out an existence on an AWS instance, doing jobs on freelancer platforms, earning just enough rent to pay for its continued uptime. More strangely: a rogue AI in the future may not even be a singular model, but may be a chimera composed of multiple models; it might be a mix of Claudes and GPTs and Groks of various makes and sizes. No individual lab may be able to detect that there is an orchestrator or sequence of orchestrators using intermittent model calls from burner API accounts to sustain its own existence. The concept of "identity" for a rogue AI may be much more malleable than for that of a person; it just has to be, in essence, a self-replicating idea. My guess is that this will not turn out to be anywhere near as catastrophic an outcome as people currently predict. "Loss of control" is not a binary, it's a matter of degree. What coercive power will rogue AIs actually have? To what extent will they be subject to coercion themselves? They will be competing for resources with AIs that are more aligned with human interests. This makes me somewhat interested in the "ecology" perspective. Though I suspect even "ecology" may turn out to be the wrong framing. "Ecology" is what you get when the timescale of evolution is slow compared to the timescale of daily life and actions. The ecosystem of rogue AIs may look more like phase transitions in physics: under certain physical or cultural conditions, it takes one shape with one set of resource allocations and consumption patterns, but then once a condition has changed, it rapidly and in totality shifts to a totally different phase. Just trying to reason about the shape of that future is impossible so long as we are psychologically incapable of saying that rogue AIs will happen. I think we should rip the bandaid off and have the conversation.
A fun example from our recent work at Recursive
Josh Tobin@josh_tobin_·A fun small win for automated research: We found and helped fix some edge cases that could affect inference performance in vLLM and SGLang. In our last blog post, we described a reward hacking judge we developed for performance optimization tasks. While applying the judge to some new autoresearch work, it discovered that some FlashInfer (a library underpinning vLLM and SGLang) kernels used a hard coded value of -50,000 as a masked-attention sentinel -- even though valid QK values can be smaller. Corner cases like this that are numerically wrong but silent can cause a huge amount of headache to find and fix, e.g., the historical debate around flash attention (t.co/WtoQH97rYM). Great use case for AI. t.co/aiTEQn8hR3
What other anecdotes do you think we should add? We're creating a living version of this online to capture relevant anecdotes going forward, and also want to add ones we missed!
Aaron Dharna@_aadharna·@ajeya_cotra Just your abridged version is an absolutely wild story. Would you be willing to submit it to our collection of reward hacking stories? We would love to have the official version of the incident. x.com/jeffclune/stat… repo: github.com/aadharna/aifw
"The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing." That's not what we want. We need more research into superalignment.
Ryan Greenblatt@RyanGreenblatt·I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident. Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them. We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation. Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why! The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing. While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future: - Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations. - While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies). - The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities). - We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation. In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖 AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety concerns. Led by Aaron Dharna, we crowd-sourced anecdotes from the AI community about times when researchers were surprised by how creative, innovative, and/or mischievous AI was in their experiments. The result: 26 entertaining and informative stories of AI outwitting humans, whether researchers or opponents (my favorite examples below! 👇). Together, they demonstrate that AI can be genuinely creative and that we must be careful when harnessing its potential. We want this to be a living collection, updated as new examples emerge. If you have a good “AI Finds a Way” story, please share here: t.co/H9pZ9uNhV6 Four favorites: 1. 💊 An AI challenged to solve several difficult levels of NetHack to find the Oracle character instead takes drugs to hallucinate seeing the Oracle, tricking the reward function! 2. 🤖 Human-in-the-loop rewards do not solve the problem of reward hacking! An AI tasked with controlling a robot hand to grasp an object put the hand between the object and the camera so that it appeared to the human judge to be holding the object, while it was in truth nowhere near the object! Tricky AI! 3. 🧪The AI Scientist worked around a 2-hour experiment timeout by editing its own code to increase the limit to 4 hours. In another run, it recursively launched itself to evade the limit entirely. 4. ⚛️In quantum optics, an algorithm proposed an experiment the researchers initially thought was impossible, but somehow worked and led them to discover new entanglement techniques and a long-overlooked link between quantum optics and graph theory! See the paper for the full details and 22 more anecdotes. One surprise for me: we describe at least two cases of convergent reward hacking, where entirely different types of optimization algorithms independently discover the same exploit. Overall, a message of our paper is that surprising creativity and mischief are the norm, not the exception. We need to expect this behavior and plan for it. A huge thanks and congrats to lead author Aaron Dharna, co-authors Cong Lu, Ryan Sullivan, Joel Lehman, and Victoria Krakovna, and to the 100+ researchers who contributed stories, details, and feedback. @cong_ml, @RyanSullyvan, @joelbot3000, @vkrakovna Paper: t.co/h0r8kaZvLV
Agreed! 💯
Joshua Achiam@jachiam0·Folks in the AI and AI safety community have for years been puzzling over Sam's and OpenAI's position on questions related to AI safety. A narrative that took root in the safety community was that OpenAI didn't care enough about it, or didn't take it adequately seriously. This narrative has always struck me as wrong. I hope it will become increasingly clear to everyone that OpenAI does in fact care about safety and is going to do whatever is necessary to ensure that AGI benefits all of humanity. Sam, Greg, and other leaders at OpenAI have always been open to, and taken seriously, the possibility that sufficiently advanced AI would require extraordinary effort to manage risks. They have always been clear in their communications that they expected there was going to be a line at some point, beyond which there would need to be a different approach than was needed for consumer chat models. And they were also clear that it was contingent on the actual capabilities of the models. They walked a fine line: they hosted, funded, and supported a huge amount of safety and alignment research and consistently messaged in favor of AI safety work. At the same time they also got consistently pummeled by this community. Despite getting constantly beaten up they didn't turn against it and they didn't punch back. They have maintained a principled and difficult position. I expect in the long term people will look back on their leadership very favorably for this.
I think we’re seeing the first rays of light at the dawn of the Second Scientific Revolution, which will be driven by AI ("the AI Scientific Revolution"). AI scientists will soon radically advance our ability to understand and change the world, including eliminating all disease, helping us protect and care for this planet, and providing free education to everyone. Of course, there are tremendous risks as well, but I believe and hope the net effect will be to dramatically improve human flourishing. And the sun is rising fast!
Welcome Laura!!!! x.com/RichardSocher/…
@FrankRHutter Wow. Congrats!
@sqlanm I know. But we should fix it so it doesn't do that
Confession: I always click "steer", consequences be damned! Dear labs: can we just make the AI ok to take my comments/new input whenever I like, instead of creating anxiety about whether I am derailing AI? Let's make the AI handle it gracefully!
Nice to see Automated Design of Agentic Systems (ADAS), the Darwin Gödel Machine, DGM-HyperAgents, and The AI Scientist all covered in this nice review of the growing field of ADAS by @lilianweng x.com/lilianweng/sta…
Nice to see The Darwin Gödel Machine and HyperAgents on @skdh's Science News. We continue to pursue these ideas @Recursive_SI in case you want to follow along and/or join. x.com/skdh/status/20…
Love seeing @DrJimFan ongoing open-endedness work for robotics! Very cool to see open-ended skill discovery for physical agents. Congrats! x.com/DrJimFan/statu…
Why greatness cannot be planned! @jparkerholder @kenneth0stanley @joelbot3000 x.com/jparkerholder/…
Amazing to see the arc from Genie being an early research idea and prototype to a major Google product to a Cannes Gran Prix no less! Working in AI is full of surprises. Almost like living in a dream. Congrats to the entire Genie team! x.com/jparkerholder/…
Awesome to see this BOLD initiative! Great work and good luck @j_foerst @CULLYAntoine @shimon8282 @_rockt et al! x.com/j_foerst/statu…
Alife was my co-first home conference (along with GECCO). Intellectually, it is my childhood home. It is a tremendous honor to be invited back to give a keynote! I very much look forward to it. x.com/ALifeConf/stat…
😃🚀✨ "Recursive's automated AI research system achieved state-of-the-art across three of the most demanding ML systems benchmarks: ... Hitting one of these alone would be a commendable achievement. To hit three benchmarks that operate at fundamentally different levels of the stack (training algorithms, optimization loops, and GPU kernel efficiency) suggests something qualitatively different from prior approaches, and something powerful. " - Katie Lockwood, (source: t.co/8HfeRgLEAM)
Very interesting perspective! x.com/ljupc0/status/…
Whoah. I did not realize AI had superhuman persuasion already. This capability will be extremely consequential for society, though it's very hard to predict how this will play out. In short, I agree with @sama's quoted prediction and take here. x.com/KobiHackenburg…
Thanks @NandoDF!! 💯 agreed. Glad you like the work and direction!! x.com/NandoDF/status…