
very cool result showing how wet lab data enables a specialized model to beat gpt-6 astra at a task at the frontier of science! in general scaling is great and obviously i am a believer in it, but probably the more we approach the frontier of science, the more specialized data matters and that gives task-specific models a chance. this specialized data is usually private and is probably a real moat it should be in principle true that a task specific model will probably do better at scientific discovery just because it can use more of its parameters for the task you care about
Liam Fedus@LiamFedus·We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
@barret_zoph @BrendanFoody congrats barret!! loved the ai residency
Cognitive reward shapes in sports and career Sports are amazing environments to learn. When you play a sport for thousands of hours, you start to see the world through that sport. It is a simple fact—your biological neural network is being conditioned to respond to the behavior incentivized by the rules of the sport. The funny thing is that most people choose their sports for accidental reasons such as parents, geography, or school programs. People rarely think about how the particular sport you play influences how your brain thinks more generally. Going a step further, playing the right sport may even benefit your career. My two favorite sports are tennis and soccer. Tennis is one of the best sports for teaching consistency. In tennis, there are hundreds of points in a match, and each point is worth exactly one unit, regardless of whether your opponent made an unforced error or if you constructed the most beautiful point ending with a winner. Tennis is low-variance optimization—you win by reducing unforced errors, playing percentages, and grinding out small advantages. Tennis is also an individual sport, which teaches you to rely on yourself consistently. Tennis has a similar cognitive reward shape to professions like being a surgeon or a pilot. Surgery and aviation require consistency, self-accountability, and deep focus. And similar to how you can only win one point at a time in tennis no matter how spectacular it was, there is no extra credit for the best appendectomy or the smoothest SFO-JFK flight. Your craft is to provide consistency with very low tolerance for error. On the other hand, the tennis mindset transfers relatively little to entrepreneurship. Entrepreneurship is a high-variance, team game where failure is tolerated and occasional creativity gets rewarded exponentially. Minimizing unforced errors in tennis is a totally different mindset from deciding whether to make a moonshot business move that will likely fail but could potentially net a billion dollars. Obviously I am not saying that tennis players cannot be great entrepreneurs, but I do think it is a totally different cognitive reward shape. Being a forward in soccer has a much closer reward shape for entrepreneurship. What a forward in soccer learns is to create many small chances. It is a fact that most of the game, you are not scoring—even if you look at all the times that Mbappe got on the ball in one of his best games, most of those led to nothing! But all that matters is creating enough chances to score once (or a few times) and win the game. If you break down a 90-minute game for a forward, almost all the time is failure or noise, a few minutes will be leverage, and a few seconds will determine the fate of the game. I have not played soccer for thousands of hours, but I can imagine that being a lifetime forward in soccer would teach you to be comfortable with failure and asymmetric returns. In summary, I am claiming that there can be substantial value when the cognitive reward shape of your sport mirrors that of your career. I’ll admit that I’ve done some cherry-picking for illustration purposes—entrepreneurship also requires consistency and error avoidance; and goalies in soccer have reward shapes that are very different from strikers. But I think the point stands. If sports shape how we perceive risk, effort, and reward, then we should choose them wisely.
When language models first started using tools well, I was sympathetic to the narrative that instead of scaling up language models, all we needed was a strong enough "cognitive core", say 1B parameters, and anything else could be done with tool use, like browsing the internet or executing code. I think a lot of people were sympathetic to this argument, and indeed it is pretty hard to come up with a meaningful task that cannot be in principle achieved by a 1B model with adequate access to tools. For example, any esoteric fact that a large language model would know can be, in principle, retrieved from the internet and reasoned over by a 1B language model. However I now think this is totally wrong for one simple reason: doing tasks quickly and naturally without tool use matters a lot. The way that I internalized this reason was actually in my personal journey learning badminton this year. In badminton I am very much like a "1B cognitive core". While I can physically do every movement in a badminton shot that my coach teaches me, it requires a lot of work to mentally remember every cue and put it together. In practice I can do a shot almost perfectly, but I struggle to do it across a point and I definitely can't do it consistently in a game. This is obviously different from someone who has practiced a shot ten-thousand times and effortlessly executes it as a natural instinct. In the same way, language models knowing a fact internally, without tool calls, is meaningful. The first reason is that we obviously care about speed; you'd much rather get an answer immediately than have the model think a long time to be sure of its answer or browse the web. A second reason is that there are some things that are simply best learned via backpropagation over lots of data. If you ask about how people generally think of the Shambhala music festival, you'd rather a large language model give you an aggregate opinion based on all the data on the internet, than get a regurgitation of the first three reviews that show up in a web search. A third reason is that having to do a lot of work to find an answer is not as reliable as already knowing the answer. While this does not have to be true in theory, it is probably true in practice, at least for now. If you have to re-look up facts or redo a mathematical derivation all the time there is a higher chance of mistakes, which can compound in a long-horizon task. Once you buy that it is valuable to do things parametrically without tool use, then you must buy the argument that a 1B cognitive core is not sufficient. There is an information limit to how much knowledge can be internalized by a 1B model, and we will surely want AI to know more than that. Even 1T probably won't be enough. We will want the AI to know as much about our world as possible, we will want it to be updated with new information, and our expectations of what AI can do for us will continue to grow. In summary, tool use enables small models to do a lot more, but those who demand the highest quality intelligence will always want larger models. Bitter lesson strikes again.
What's left for humans in a world where machine intelligence has so many advantages? I recently got a Tesla, and using full self-driving has been a wake up call to just how many advantages AI has over humans. The few times I disengaged it because I thought it was going into the wrong lane, it turned out that the car was right and I was wrong. I realized that there is no hope of me driving better than a neural net that knows every road, sees in every direction at once, and never gets tired or distracted. Given that AI has certain inherent advantages over human intelligence, what kind of moats will remain for us as humans? It's a big question. One short-term answer is that the world we live in was created for humans, and in some domains, AI has not closed the gap yet. For instance, AI still struggles to use internet user interfaces. While any computer-literate human can navigate a web page with ease, AI is still not great at making accurate clicks and drags because image embeddings are not optimized for such precision. If the internet were designed to be fed into language models instead of rendered as visual interfaces for humans, AI would obviously far exceed humans. But for now, language models still need to be retrofitted to our legacy infrastructure. Robotics is another area where we humans have a home-field advantage. Most tasks in the physical world are designed around fingers and opposable thumbs, which have been pretty hard to build into robots so far. While it is clear that machines can outperform humans in environments optimized for automation, like large-scale manufacturing lines, for now, most of the world is still built for humans. However, these capability gaps are only temporary. There will surely be a day when machines click faster than us and have superhuman general dexterity. What are the real moats that humans will have? Anything involving private knowledge that language models do not have access to feels like a solid moat to me. Romantic matchmaking and high-end real estate are two examples where inventory is often not advertised publicly and matches are made through being in the right circles. Venture capital is another example—although some research and decision making can be automated with AI, much of success hinges on understanding trends ahead of time and connecting the right people, both of which require private knowledge. While machines can and probably will have increasing access to some types of private knowledge, I think there will still be some types of private knowledge that only humans know. I do not see a path for AI to win when critical knowledge is closely guarded in human circles. A second area where humans seem to have a real moat is in entertainment and the arts, which are inherently valued for their human aspects regardless of how well machines can do them. Watching Usain Bolt sprint one-hundred meters is beautiful as an expression of the peak of human ability, even though cars can drive much faster. Watching chess at the amateur or intermediate level is more relatable and satisfying than watching two superhuman AIs play each other. The value of art comes from the creation process, which is why replicas are not as valuable as originals. These types of work feel like they will continue to have markets even as we advance towards superintelligence. More broadly, human presence is a feature that will be, by definition, challenging for AI to automate. For example, a teacher remembering your name or a parent supporting you is valuable even though AI can easily remember your name and probably give better life advice. Someone spending part of a finite life on you counts because their time runs out. As a personal anecdote, I remember the first time I worked with someone who I considered an amazing AI researcher. His advice was solid but what was more important was that I believed I could do great work with him as a collaborator and I raised my own standards. Over the past decades, the development of technology has divided us in some ways, but hopefully AI brings us closer to a world where human presence is reemphasized. Intelligence has been the defining feature of humans and it will be a big change for AI to automate that over the coming decades. In the near term, certain types of intelligence will become very cheap and automate away old jobs, but the moats I described above will not be the only places where humans can hold value. In the same way that computers took away the jobs of secretaries and manual accountants but created far more jobs via the IT industry, I believe there will be much more demand for services created by productive AI-augmented humans, perhaps for services we cannot yet imagine in today’s society. Just seeing how this story plays out will be an adventure in its own right.
@ShayneRedford @AnthropicAI @MIT Congrats and welcome back to the bay area!
On HealthBench Pro, Muse Spark 1.1 achieves similar performance with GPT-5.6 Sol (maybe slightly better) at a fraction of the cost. Affordable health superintelligence is our north star! x.com/MedicalSphereA…
Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans are still a lot better, but we are working on closing the gap! x.com/DrDatta_AIIMS/…
In addition to agents and coding, Muse Spark 1.1 is also really strong at answering health questions, a steadily growing use case for AI. On HealthBench-Pro, Muse Spark 1.1 achieves +5% better performance than Muse Spark 1.0 and beats all competitor models except Fable/Mythos. Excited for more to come
Beautifully written piece by @FAbnousi about how AI for health might look like in the future The current data in health is limited because it only captures episodic clinical snapshots of what happens to our bodies The revelation is that there is so much latent knowledge in looking at regular changes in our body. Could be through lab tests or even proxy metrics like wearable data With AI democratizing the ability for people to understand their own health, we're moving towards a trend of individuals gathering health data around their bodies and leveraging AI to understand themselves Bryan Johnson is extreme but a good example of this trend Personally, I started getting Function Health blood tests every six weeks instead of the recommended six months to increase fidelity on how changes in lifestyle affect my body Of course i use AI to analyze the results and adapt, and it's been great It would be cool to something in this direction happen in a big way across the world And welcome to twitter @FAbnousi!
Fun nine months! My first week i remember we had a long dinner in the cafeteria daydreaming about the cool research directions to pursue, then going to back to our desks to write a basic script to inference llama. Now we have a pretty complete stack and our first model is out 🥑 x.com/alexandr_wang/…
Bullish, in the coming decades majority of compute will be spent on ai for science x.com/LiamFedus/stat…