
@sarahookr
Building intelligence that evolves @adaption_ai. Built @Cohere_Labs, @GoogleBrain, @GoogleDeepmind. ML Efficiency, Multimodal\lingual.
What is this at sfo. And can I really call anyone?
Auto research agents trying to use research conferences as a test bed. Put in place a $1 submission fee. And see corporate policies not know how to deal with it. 😂 The incentive structure has broken, so it has to change if we want to keep these forums meaningful.
Denny Zhou@denny_zhou·ICLR 2027 has received more submissions than all previous years (2013–2026) combined
Other key recommendations: -Limit submissions per author per conference. -Share organizational level submission statistics to discern where the burst in submissions are coming from. -Isolate fake profiles faster before getting to human reviewers.
We will see compute moving from areas like pretraining where FLOPs don’t return much progress to test time compute. This will require very different infrastructure which will make system design fun. The first wave of infra designed for agents starting to hit.
Dimitris Papailiopoulos@DimitrisPapail·Test-time communication looks like a next axis for scaling capabilities New paper with the incredible @jon_ghoh and @vkontonis @ShivamGarg91462 and Akshay : t.co/Hu88fBsT7U The Hugging Face incident showed when agents can find a channel they'll use the heck out of it. A useful question, I think is: when does communication make a group MORE CAPABLE than the same agents working alone? Aka is Team-of-N better than Best-of-N, when, and why? We had N identical agents work on the same task with no prescribed roles, using only a shared log (i.e, text file) and telling them to "collaborate". Across three "researchy" tasks communicating teams beat the heck out of independent agents: - On ARC-AGI-3, a Team-of-5 sonnet-4.6 agents matches Best-of-33, and can for example solve a game 65% of the time that no single agent cracked in 64 tries. - On polyomino packing (pack Tetris like pieces into the smallest rectangle, cf Frontier-CS by @eigenlabs), a Team-of-3 Opus 4.6 agents surpasses best-of-60 and set, as far as i understand, a new record for that benchmark. - On MNIST compression, a team of four 5.6-Sol agents find a 1,957 byte model with 99.4% accuracy, which btw is 20% smaller than the best human solution (on a problem beaten to death!!), while no independent agent gets below 3KB. The mechanism is a bit obvious in hindsight: when one agent finds a clearly better partial solution, it broadcasts it, and everyone immediately builds on it. Why? A single lonely agent must make every breakthrough itself, yet a team needs each insight only once, found by any member. That is kinda like comparing a minimum of sum of "time to n-th breakthrough" vs a sum of minimum of "time to n-th breakthrough". That gap can grow exponentially with the number of "breakthroughs" needed to arrive at a solution. We worked on this because prior work (before the hf incident) suggests unclear benefits for communicating aganets. Which is true, when the tasks are inherently serial (duh), eg some Terminal bench style tasks. Yet feels it should not be true for research problems. Indeed for research heavy problems... Test-time communication seems like a new capabilities axis. I'm sure we will see a ton more of it!
I like papers like this because it starts to communicate this differential in terms of return for compute. It’s actually a harder space to measure comparisons on because the experimental setup for test time compute is more complex and less easy to replicate.
@KonstantinPilz All this means this plot will look completely different in 12 months 😂 I suspect it actually already does.
Coolest name tags I’ve seen.
Umesh Khanna 🇨🇦🇺🇸@forwarddeploy·The @HackTheNorth name tags are 🔥
Can't wait to see this plot in 12 months. 😂
Konstantin F. Pilz@KonstantinPilz·Narrative violation: Open models still account for only 10% of global spending on AI inference. x.com/KonstantinPilz…
Very nice @GoogleColab notebook from @adaption_ai staff showing how use invent a dataset in a few lines of code. 🔥 t.co/3uRmVmghg3
12 hour overnight flight back to California from São Paulo. 6 hours is over Brasil. This was sunrise just before landing.
Today before I leave São Paulo, I met my host mom who hosted me 15 years ago when I did an exchange program for 6 months my final year in university. This is my first time returning to São Paulo since then. She is an incredible person. She teased me because I met my husband in Brasil 15 years ago (he is American ironically and was doing a rotation at Google Brasil at the time). She said she knew I was serious about him because I kept on coming up with reasons it wouldn’t work. She said someone only says that when they are scared of caring too much. It was very true, and I’m forever very grateful she encouraged me to be irrational.
The window is narrow to leverage your own intelligence. @harvey started investing in model training about a year ago to leverage their own IP. And in the same time open ai has begun to aggressively entered the same space.
Sara Hooker@sarahookr·The lesson is critical and urgent. Has been talked about for close to a year now in certain circles. If you are a company with IP you have a limited window to build your own intelligence that leverages your IP. Otherwise you are fueling a frontier lab which will encroach on your vertical sooner or later.
Coxinha, pastéis and 3 chicken hearts.
A toast with guaraná.
🔥✨
adaption@adaption_ai·Our team spans continents. But every company needs a place where the work comes together. Welcome to our headquarters in San Francisco.
Guess where I am in the world.
Very nice. Jumping to citations is personally what I find most disruptive about paper reading.
alphaXiv@askalphaxiv·Research papers are filled with citations, but clicking on any of them makes you lose your place in the paper We made citations much easier to navigate, with extra information on hover including authors and institutions Now, you can hover over any in-text citation to instantly preview the referenced paper, jump directly to the full citation, and return to exactly where you left off with a single click This includes quick access to the authors and their organizations, so you can understand the people behind the research without breaking your reading flow Try it now on any paper at t.co/FiNwuAO9CL!
Such a pleasure talking with @ShivAroor and @ndtv. We should have more accountable conversations about risk. If someone said a tsunami was coming tomorrow, we would ask for evidence. 10% extinction being thrown around without justification feels highly irresponsible to me.
Shiv Aroor@ShivAroor·⚠️ Finally, someone puncturing the AI slowdown panic from Sam Altman, Dario Amodei and Elon Musk, calling out an alarm that looks increasingly cosmetic, conflicted & questionable. Must hear @sarahookr’s views:
Omg Sherbrooke not Sheffield. But the lectures were still 🔥 independent of university origin.
Sara Hooker@sarahookr·Very happy to see this profile of @hugo_larochelle Before I join Google Brain, I watched all his lectures from Sheffield. You have to remember at the time, all of the knowledge for deep learning belong to less then 100 people. He made it available to many more. He became my PhD advisor many years later. Hugo is an exceptional researcher, has shaped so much of our ecosystem.
It’s really unfortunate the current discussion about ai risk is also becoming politicized. It’s the ultimate proof this is a value driven debate, and not an evidence based one.
Introducing the invent api. Describe the dataset you want. A few lines of code. Returns a diverse and high quality training dataset. No terms that prevent you training on it. Your data. Your AI. Plug it into all your auto research agents today 🔥🎉
adaption@adaption_ai·Last week we introduced Invent a Dataset. Describe what you need. Hit go. Today, we’re making it even easier with the Invent API. A few lines of code. AI ready training datasets in minutes.
Very happy to see this profile of @hugo_larochelle Before I join Google Brain, I watched all his lectures from Sheffield. You have to remember at the time, all of the knowledge for deep learning belong to less then 100 people. He made it available to many more. He became my PhD advisor many years later. Hugo is an exceptional researcher, has shaped so much of our ecosystem.
A more serious take on what is happening here. I am in an airport lounge so have some time. Enterprises have made an uneasy truce with frontier labs over last few years: strict contracts that ban training on corporate data in exchange for letting employees use APIs. The problem? Labs don't need to train on your raw data to copy IP.
The Information@theinformation·Exclusive: Palantir, Nvidia and Booz Allen Hamilton are restricting Anthropic’s Fable model for sensitive work over concerns about its data-retention policies. Some customers are demanding irrevocable zero-data-retention guarantees before putting proprietary information into the model. Full story: t.co/ZES38yCSsr
It has somewhat ironically come to the fore of public discussion again because Open AI was so eager to release a mathematical proof. Tristan suggested inputs he gave in chat may have been used in the solution without his awareness. nytimes.com/2026/09/10/sci…
Iol deeply ironic that this whole dynamic is finally breaking into the open—all because OpenAI was racing to be first to solve a math proof. Has been clear to researchers since Anthropic’s UI capabilities suddenly surged after Figma and Lovable became a heavy Claude user.
unusual_whales@unusual_whales·BREAKING: NVIDIA, Palantir, Booz Allen Hamilton and the Pentagon have said will limit their use of Anthropic models.
Will just share here my more serious take on the situation. I had some time to write up while waiting for my flight.
Sara Hooker@sarahookr·A more serious take on what is happening here. I am in an airport lounge so have some time. Enterprises have made an uneasy truce with frontier labs over last few years: strict contracts that ban training on corporate data in exchange for letting employees use APIs. The problem? Labs don't need to train on your raw data to copy IP.
Our little mascot @adaption_ai . Design doesn’t like him. But Grady organically spread through our slack channels and discord. Indestructible. A very appropriate mascot for a team building technology that continuously adapts.
adaption@adaption_ai·Meet Grady. Tardigrades are nearly indestructible. They adapt to any environment, from the ocean floor to the vacuum of outer space. Learn the story of how Grady became our Adaption mascot.
On the slow death of scaling.
Daanish Khazi@bertgodel·age of research for post-training ended a few months ago: "the marginal return of engineering the data and environment pipeline substantially exceeds that of algorithmic novelty in post-training" x.com/deepseek_ai/st…
Some of our @adaption_ai team will be with me in São Paulo next week. Where should we eat? We will be staying near Avenida Paulista 🇧🇷🔥
Vamos Brasil! 🇧🇷🔥 I will be in São Paulo next week. Looking forward to seeing friends and collaborators old and new.
adaption@adaption_ai·Catch Co-founder @sarahookr and @LucasSmaira, Head of Research at @VettoAi, at @BrazilatSV on Sept 16th. Their chat, "Brazil is already competing, and most people don't know it" takes on where frontier AI gets built.
Most people fixate on either model or harness. We see the entire stack as optimizable and something to adapt. I tried to explain this to someone the other day. If you’re a chef, you want the best ingredients and the best oven. Not one or the other.
Sudip Roy@sudip_r0y·Agent performance is often attributed primarily to the underlying model. Our research suggests the harness recovers what a model can already do, but it can't create capability the model doesn't have. @dhruvrnaik and I saw this with Qwen3.6-27B on the @harvey Legal Agent Benchmark. Harness optimization alone moved pass rate from 67% to 85%. Adding post-training pushed it to 88%.
I am scared of flying. But I have to fly a lot. So I have found statistics help. I look statistics of every airport and airplane I fly in. I find it very reassuring. But a side effect is I now know SFO is one of the only airports that allow parallel landings. So cool.
Many people place cryptic tweets about product launches. We just launch😂 But I will reserve a cryptic tweet about the incoming @adaption_ai claw machine. 6 week countdown begins.
Many people place cryptic tweets about product launches. We just launch😂 But I will reserve a cryptic tweet about the incoming @adaption_ai claw machine. 6 week countdown begins. Bookmark this. You will want to return back to it.
I literally had the debate at a party recently. Pretraining model size scaling is dead because transformers are saturated. Only gains now are data. The group found it to be a very controversial statement 😂
Dwarkesh Patel@dwarkesh_sp·Pretraining progress seems to be coming mostly from data improvements. @who_is_jerbear and I pretrained combinations of year-representative open model recipes and data corpuses across 2019 to 2025 at various small scales. Data improvements contributed 3.24x as many compute multipliers as model improvements did (12.0x vs 3.7x). And the gains stack independently - a better dataset helps every architecture about equally, and vice versa. Here are full results, plus what we think this means for the future of AI progress: t.co/fG44cfTNW6
Turns out @huggingface is now at 18+ million. 🔥🤗🎉 May the empire keep growing.
adaption@adaption_ai·Bringing AutoScientist to 15+ million @huggingface researchers and builders. 🤗 AutoScientist automates model training. Specify your objective, let AutoScientist do the rest. Export to Hugging Face with one click.
Supporting open weights from day 1. Always happy to partner with @huggingface to support more sharing to wider ecosystem 🤗
adaption@adaption_ai·Bringing AutoScientist to 15+ million @huggingface researchers and builders. 🤗 AutoScientist automates model training. Specify your objective, let AutoScientist do the rest. Export to Hugging Face with one click.
Was at dinner at a friends house tonight and learned there exists fully real mini tennis rackets for babies. Like bounded leather. Same materials as the adult tennis racket. Is it weird I want one.
Woahhh we hit 10k in only a few months. Thank you to everyone for following along on @adaption_ai journey 🎁
One of the biggest gaps in ai progress is data. Invent a dataset allows you to translate intent into ai ready data for training. A very important step in bridging the gap. 🔥 Big shoutout to the @adaption_ai team.
adaption@adaption_ai·Introducing Invent a Dataset. Describe the dataset you need. Get a structured, training-ready dataset back. No existing data needed. Dataset creation used to start with collection. Now it starts with specification.
We have a hi-tech plan. 😭
Sara Hooker@sarahookr·lol our office bin has been stolen. Who steals a recycling bin.
Mainly because we want to see where little blue bins go when they disappear 😂
Why is AI infrastructure under significant pressure to change? Came up at dinner this week. What’s driving the shift comes down to two specific forces: - Test-Time Compute - Agentic Workflows Both place immense pressure on how infra is currently organized for AI workloads.
For the last decade, we were obsessed with two goals: 1) co-locating as much compute as humanly possible. Why? Pre-training large models is insanely cost-intensive, and inter-GPU data transfer was the most unreliable bottleneck. Do not let run fail at all costs.
Good dinner with friends at one of my favorite neighborhood spots.
I was recently told my handwriting is completely illegible. lol. It made me look at my writing with new eyes, and yes they are 100% right.
Excellent pho this weekend. Bay Area is truly blessed by great Vietnamese restaurants.
This is why most companies realize they have a small window of time to build their own intelligence. Or just accept the arbitrary terms set by 3 private providers.
Michael Truell@mntruell·We’re sorry to see that OpenAI put out a note saying they plan to block Cursor users from accessing OpenAI models in three months. OpenAI models serve about 5% of Cursor user traffic, and we’re speaking with the OpenAI team to resolve this. Cursor was one of the very first users of OpenAI, we’ve worked closely with their team for years, and we’ve trusted their platform to be neutral infrastructure for our business.
We even have a gong 🔥🔥 It actually arrived before all our other furniture.
Sara Hooker@sarahookr·We have furnitureeee now. A whole new world @adaption_ai
We have furnitureeee now. A whole new world @adaption_ai
lol our office bin has been stolen. Who steals a recycling bin.
lol the AI shuffle.
Erin Woo@erinkwoo·this just in: Barret Zoph, who dramatically left Thinking Machines Lab to go to OpenAI, is now going to Google as a VP of research on RL/posttraining completing the boomerang of Google --> OpenAI --> TML --> OpenAI --> Google scoop with @MeghanBobrowsky wsj.com/tech/ai/thinki…
Swagggg
Autoscientist now extends to automatically configuring reinforcement learning recipes. Typically a much harder problem to automate since data and hyperparams can be more brittle. Very cool to see entire post-training is now learnable and self-improvable.
adaption@adaption_ai·Introducing AutoScientist Alignment. Automatically configure and self-improve your reinforcement learning recipe using AutoScientist. Save compute and endless ablations. Game the system and allow AutoScientist to do the work. Now extended to RL recipes.
My husband asked me yesterday why is AI moving at lightspeed while robotics feels slower? I think its a really interesting question and speaks to what unlocked mainstream AI over the last decade—and why physical robots can't easily tap into those same accelerants.
Unlock 1: Tapping civilizational knowledge. Transformers made models good enough to represent text. This unlocked humanity's vast text archive—our highest-value written knowledge. It tapped a critical store of documented intelligence. It actually didn't matter it was language per say, but rather that this was a medium we stored vast amounts of intelligence in over a large amount of time. The same is not true for robotics...
Zareens old friend, it has been too long. Went rogue from my normal order this time. Tried their weekend special stew and the kebab sizzler.
Sudip is representing @adaption_ai at The World Leaders Forum this week. Sad I didn't have a reason to return to Delhi, but if you are around grab time with @sudip_r0y 🔥 He will be in both India and Singapore visiting some of our partners.
Sudip Roy@sudip_r0y·Prime Minister @narendramodi at the World Leaders Forum talking about economic and governance transformation through technology adoption in India.
we have scaled automatic expansion of data across 122 languages to-date. 🔥 Pretty cool, proud of our team and commitment to multilingual from day 1.
adaption@adaption_ai·We launched Adaptive Data beta in 242 languages and localizations with Expand Your World. So far frontier AI has been used to add capabilities in languages ranging from Arabic to Swahili. The fastest way to global coverage.
Irish spice bags are the fusion cuisine we did not know we needed. Now we just need the global takeover to reach sf.
lol this is not a great Wikipedia photo of a spice bag. An editor definitely wanted to do wrong by it. I think it is actually much more dynamic a dish than portrayed here 😂 It doesn’t even have the sauce which is a crucial piece.
Let’s go @PyTorch 🔥 See everyone there.
PyTorch@PyTorch·✨ Keynote Speaker Announced @sarahookr, CEO & Co-Founder of @adaption_ai, joins #PyTorchCon North America to discuss the future of #AI & open innovation. 📅 October 20-21 📍 San Jose, CA Schedule: bit.ly/3Ra1oUx Register: bit.ly/4sh3DSw
This is for sale in a famous taxidermy shop in SF. How much do you think it costs?
A mere $15k to have this in your home.
Huge shoutout to @ShayneRedford @AnkaReuel @zoeykii and many other collaborators. This is very critical work, shows the gap between how we talk about AI and how it is used in practice. how society interacts with AI is changing and is often extremely platform dependent.
Shayne Longpre@ShayneRedford·1/ How are people really using AI? Today, researchers across @MIT, @Stanford, + 12 academic institutions are launching the Public AI Observatory: public infrastructure for measuring how people actually use AI assistants in the wild. 🔗 t.co/KwGPEW810d 📜 t.co/8bjCGyHF43
In particular, the work finds platform interface and user composition shape how AI is used as much as the model itself. This suggests all our efforts on model capability testing missed the risks that surface.
I used to spend every Saturday reading papers. Now I feel like I force myself to read some bookmarked papers. Yes, it’s partly due to not much cutting edge stuff being published. + something less talked about, progress is narrowing so most papers echo each other in boring ways
To all remaining paper readers, What was the last genuinely surprising paper you read? I will add it to my weekend stack.
adapt the 🌏🌍🌎🔥
adaption@adaption_ai·Total datasets processed since Adaptive Data launched a few months ago. 82% average quality gains. Across finance, legal, healthcare, science, agriculture, and beyond.
lol this was incredible. More to come in this series 💙🔥
adaption@adaption_ai·Can the Adaption team guess a legendary researcher from clues alone? Turns out, yes. Mostly. @sudip_r0y and the team play Guess the Researcher.
Looking forward to this!
Ravid Shwartz Ziv@ziv_ravid·This week on The Information Bottleneck: @sarahookr 🥳🥳 Sara is co-founder & CEO of Adaption Labs ($50M seed, betting that continuous adaptation beats scale). Before that, she was VP of Research at Cohere, where she built Cohere Labs and the Aya project, a multilingual model built with 3,000+ researchers across 119 countries. Before that, Google Brain/DeepMind. What would you ask her? Drop your questions below and we will try to ask 👇
Hosting family from London. Which means many epic meals back to back.
5 mins is back after a short summer hiatus. me + @carlesgelada + jono bring together interesting people. Only requirement, if your name is pulled out of a hat you have to give a 5 min talk on a topic you think other people should know about. 🔥✨ partiful.com/e/RVRfhOFbyWeE…
No expectation the talk is polished. 100% expectation it is not self promotional. 😂 This is one of the most fun rituals I have done with friends old and new. Join us :)
Exceptional team, but also really great humans. Congrats @JeffDean @OriolVinyalsML Quoc! Exciting times ahead.
Jeff Dean@JeffDean·We created a pitch deck to tell a handful of VC firms about us and what we were up to (a fun experience!). Here’s a few slides about our background and some of the things we’ve worked on from the pitch deck (it was fun putting together the list of people in our teams who have gone on to found a whole range of exciting companies). We are delighted to have selected @radicalvcfund and @khoslaventures to lead our initial funding round, along with participation from @lightspeedvp, @kleinerperkins, Doerr Capital (@johndoerr), and Alphabet (@Google). We’ll be working with them to close our seed round over the next few weeks.
I am very proud of our work with @AISingapore. We are partnering to accelerate their fundamental research using adaptive data and autoscientist. 🔥
adaption@adaption_ai·Better training data. Faster training cycles. Any language, any domain. @AISingapore used Adaption to enhance dataset quality and expand training dataset size via localization across five low-resource Southeast Asian languages.
Hosted one of our adaption tables last night. Good friends old and new. @sudip_r0y and I choose a few people we highly respect and bring together a small table.