
@natolambert
Open model research @ something new. Prev. co-led Olmo at Ai2. Writes @interconnectsai, wrote https://t.co/alRXKINTwE
OOO for a bit, time to celebrate & rest :)
In the frantic era feeling early RSI effects, sending friends physical books, hosting events, chatting with students early on their ai journey, etc are all wonderful to do. Many things AGI won’t replace, and it’s easy to forget to do them. I’ll visit a few more universities!
Yacine Mahdid@yacinelearning·just received a great bed time reading book by the mail looking forward for tonight
Over the weekend I had Codex parse 500K arXiv AI/ML papers since ChatGPT to understand which open models are used for research. In 2024, ~30% of papers mentioned an American open model and only 10% a Chinese model. Today, ~40% of papers mention a Chinese (open) LLM, and only 25-30% an American one. Chinese models are the default for research. Chinese mentions are still growing while American open models are stagnating. When looking at this data it's important to remember that papers substantially lag model releases, as research takes a long time. Qwen's steady growth is reflective of this, but so is Llama's lasting power. Some more observations: 1. Qwen has been steadily growing, and today 1/3 of papers which mention any LLM mention qwen. OpenAI's closed models are the highest overall, at ~37%. 2. Llama peaked around April of 2025 at 30% of papers which mention any LLM (including ChatGPT etc). Llama 4 was released at about the same time, and Llama has been declining since. 3. Gemini and Claude are less common than the leading open models, mentioned in 10-15% of papers puts them behind all of Qwen, Llama, and DeepSeek. Open models should be and are the foundations of open research. The % of papers mentioning any LLM have been steadily climbing since 2023. | Year | January | April | July | October | | 2023 | 10.43% | 15.39% | 18.69% | 32.18% | | 2024 | 29.70% | 33.93% | 35.70% | 44.25% | | 2025 | 39.23% | 45.28% | 44.94% | 53.52% | | 2026 | 55.49% | 57.26% | 53.14% | TBD Now over 50% of AI papers, from 10% in 2023. Other notes: - Gemma and Mistral hover around 5-10%. - Our beloved fully-open Olmo models have been ~1% since the first release in Jan. 2024. - DeepSeek has a clear jump after R1 in Jan. 2025 - Data derived from the most popular ML arXiv categories: cs. AI, cs. CL, cs. CV, cs. LG, stat. ML Just like our downloads and derivative model data, this is updated daily on the Interconnects Open Model Dashboard.
Data is here: dashboard.interconnects.ai/?axrange=all#a…
So happy for the whole Marin team. They've been quietly working up to this for a long time. I'm excited to get to work with the community post-training this model!
Percy Liang@percyliang·🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
What's the best work you've consumed on open models, us v china AI, distillation, or related topics? Blogs, papers, videos, podcasts all welcome.
This really needs to be repeated more often. It really seems like maybe voices to organize it and instruct people trying to navigate the sea are what missing. Rather than the information itself.
Séb Krier@sebkrier·It's underrated how much you can learn about AI with what's public and out there already. People overestimate the degree to which you need to be in a lab to understand these systems and the direction of travel; don't let the status arbitrage discourage your independent research!
The first LLM lab to majorly downsize on employees ahead of their next training run
Florian Brand@xeophon·1) what x.com/Techmeme/statu…
I think this’ll be my first MATS cohort I’m mentoring! I’ve met so many wonderful MATS alumni, excited to contribute.
MATS Research@MATSprogram·🚨 MATS Winter 2027 applications are now open. Fully-funded, 12-week fellowship for aspiring & established AI alignment, interpretability, security, governance researchers & field-builders 📍 Berkeley/London 📅 Jan 19–Apr 10 💰 $6.4k/mo + $8k - 16k/mo compute Apply by Sep 6 ↓
I think the RAM score @xeophon and I made for ATOM work is a really solid metric. If I just rank by the top RAM scores @ 30days within our artifacts curation, all the models are really solid. RAM score @ t= (model's cumulative downloads) @ t / (top-10 downloads for size class) @ t RAM is essentially looking at a model's downloads at specific points in time (this case t is 30 days), relative to the downloads of the top 10 models of all time in that size bracket. A score of above 1 means it's on track to be a top 10 downloaded model at that size category. Imo most meaningful for 50B+ models. Downloads alone normally isn't great on HuggingFace, but it's rare a large MoE gets a ton of adoption without being a solid model.
artifactshub.ai/?sort=ram&pick… atomproject.ai/relative-adopt…
By the way if you are a writer and not routinely freaking out about “whether you still got it” you aren’t doing it right. That anxiety is what drives you forward.
Dean W. Ball@deanwball·By the way if you are a young writer and not routinely freaking out about “whether you still got it” you aren’t doing it right. That anxiety is what drives you forward.
The amount of safety work that would be done if frontier labs were required to release a certain amount of RL rollouts for public safety checks would be incredible. We can start advocating for this once we kill the distillation mind virus.
We should have independent organizations that can access the full details of these training runs for monitoring, not just after various accidents occur. It's good that OpenAI is sharing this, but the best way for safety is more trust and more eyes on the hard problems to solve.
Sam Altman@sama·We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment. We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress. We are optimistic about the alignment work we are doing, and we remain committed to making frontier capabilities widely available. t.co/51kvKfbfrO
Claude got scooped on curing cancer, ngmi
The Kobeissi Letter@KobeissiLetter·BREAKING: Moderna stock, $MRNA, surges over +110% after announcing the first ever positive Phase 3 results for a personalized cancer vaccine.
Happy to see this! The written word is extremely powerful and we’ll learn a lot about LLMs in making them master it.
Rosmine@rosmine·Announcing Deft, a new AI lab for better writing, cofounded with @jmrphy See the picture for launch announcement the Deft model wrote for itself Currently, 86% of user queries are fully human according to pangram. This is still a small beta model and it might make mistakes. We are launching our public beta now to get more feedback before scaling up. Tips for better performance: - Add more details to your prompt. If you just provide a short sentence prompt, it will likely get detected as AI. - Try changing the style in "advanced options" - Deft currently works better for some use cases like Analysis/Essays, Creative writing, and Rewrites. It works less well for Marketing copy and news articles. Our main goal is better writing, fooling AI detectors is just a side effect. Link below
Fooling AI detectors is silly to include tho, is a total side quest.
Personal milestone: 1000 true fans of @interconnectsai ! Hitting a very long term goal feels great. I’m very happy to get to be an independent voice in AI. Cultivating a paid base helps me commit to that longer term, and scale Interconnects’ impact.
I tell a ton of people that getting a PhD is not a good idea for your career, but it’ll likely be far higher in personal growth and fulfillment which is maybe worth way more anyways. I need to think about how to price that in.
Nate Silver@NateSilver538·Rare unironic life advice. It can be great for many people but if you're on the fence about whether to go to some sort of graduate school, i.e. you don't have something fairly specific you're trying to accomplish, probably don't. A much worse value proposition than 20 years ago.
Nvidia’s optimism on open-source AI harkens back to Meta & Google investing so much in the free web with the assumption that if more people get more access to the web, good things will happen. Nvidia's betting AI will play out similarly -- massive explosion in opportunities.
Interconnects AI@interconnectsai·Teaching Everyone to Fish for Tokens Nvidia wants you building your own model, not buying from Anthropic/OpenAI. interconnects.ai/p/teaching-eve…
I hope @NVIDIAAI is as successful as google was in open source ;)
The core of open-source AI is that the training recipe itself is the closest analogue to Linux and open-source software. Model weights are transient. Nvidia is building their models in the open so everyone can build their own later, avoiding a monopoly on intelligence.
Interconnects AI@interconnectsai·Teaching Everyone to Fish for Tokens Nvidia wants you building your own model, not buying from Anthropic/OpenAI. interconnects.ai/p/teaching-eve…
IMO OLMo 1-3 still underrated in their contributions to science. I wish we could’ve spent more time amplifying this while building them.
badlands
This happened 12 months ago…
Polymarket@Polymarket·JUST IN: Alibaba’s Qwen becomes the world’s No. 1 open AI model by downloads, topping 3 billion globally.
Fixed. Thanks @huggingface team!
Nathan Lambert@natolambert·PSA: We're seeing that a few weeks ago, pretty much all models on huggingface had an ~30% sustained reduction in daily downloads, e.g. making our august prediction on the @interconnectsai dashboard meaningfully lower. @julien_c or @huggingface did you change some filtering?
Happy to see @NVIDIAAI released their expert models for MOPD. Starting to make research there much more accessible (tho these are big models for most researchers). huggingface.co/nvidia/NVIDIA-…
With prodding from @xeophon I added a crucial detail. Data industry go brrr in China. Our interconnects group chat has on many occasions been discussing the data industry in China recently, a huge change from when we visited in April.
Nathan Lambert@natolambert·GLM 5.3 notes and why we should stop being so surprised about these very strong Chinese models (most of this is talking myself through some of my denial -- yes, these models are the real deal). interconnects.ai/p/glm-53-how-c…
GLM 5.3 notes and why we should stop being so surprised about these very strong Chinese models (most of this is talking myself through some of my denial -- yes, these models are the real deal). interconnects.ai/p/glm-53-how-c…
Farewell Seattle! It's been such a wonderful life/career stage for me. Onto new adventures (and mountains).
Lots of people posting about Z ai “benchmaxxing” to make this model. I think the truth is messy and many faceted: 1. Yes Zai probably cares slightly more about public benchmarks than OpenAI/Ant, helps with marketing 2. Zai is not benchmaxxing to the point where the model is fried (if they did, they’ll fix it - they likely use this model internally a lot) 3. Zai likely has a narrower distribution of tasks the model is great at. 5.2 was great at agentic stuff, not the rest 4. GLM has not had vision - single modality is definitely easier 5. Zai is definitely extremely good at what they do, likely well more compute efficient than OpenAI/Ant 6. Time to release for Zai is likely days not months like OpenAI/Ant - this massively flatters them All together, seems like a perfectly good strategy and congrats on the release - excited for the weights to be out for more broader testing.
Z.ai@Zai_org·Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: z.ai/blog/glm-5.3
Good on Google for shipping this. Looks like a solid model, and the only way past the haters is through the pain of public feedback. Better to do it sooner than later.
koray kavukcuoglu@koraykv·Today we're launching Gemini 3.7 Flash - our latest workhorse model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. ⚡️ We have been iterating rapidly with the Flash series, going from 3.5 to 3.7 in just 3 months, making it more helpful across a wide range of tasks: • Software Engineering (DeepSWE v1.1): 37.0% ➔ 65.3% • Web Development (Code Arena Elo): 1506 ➔ 1588 • Enterprise Automation (AutomationBench): 13.4% ➔ 30.4%
Collecting open model stickers for my Spark. Thx @Zai_org and @arcee_ai
@scaling01 Mostly that the gap isn’t that giant now but will be
I made a short video covering the history of Shoggoths, as used as a post-training meme and also deeper metaphor. Wild to think the new generation of people working post-training may not even know about them 😱. Still a very apt metaphor imo.
The original :)
w̸͕͂͂a̷͔̗͐t̴̙͗e̵̬̔̕r̴̰̓̊m̵͙͖̓̽a̵̢̗̓͒r̸̲̽ķ̷͔́͝@anthrupad·The vibes shifting from "Anthropic is so far ahead" to "model competition back to all time highs" took like 4 weeks. 🎢
My latest piece is a little verbose, but I think it is really a nice example digging into a sort of knowledge work (and arguably science) that the models really aren't close to solving today -- writing something like a AI textbook in a well-known field. interconnects.ai/p/i-wrote-an-a…
PSA: We're seeing that a few weeks ago, pretty much all models on huggingface had an ~30% sustained reduction in daily downloads, e.g. making our august prediction on the @interconnectsai dashboard meaningfully lower. @julien_c or @huggingface did you change some filtering?
@interconnectsai @julien_c @huggingface dashboard.interconnects.ai
A few years ago, when I started my AI textbook, I would've guessed AI models would've maybe made it irrelevant at the time of publishing (now in 2026) due to them getting way better at a fairly simple task. I was surprised to be wrong. Today, models haven't gotten much better at non-fiction/technical writing in long-form. This has me worried for AI models' abilities to do genuine open-ended science, one of the core posited benefits of models. Yes, the models can solve known math problems etc. and make awesome connections across fields, but how are they going to build a literature themselves without massively increasing entropy? Being able to write and organize thoughts is crucial to advancing collective knowledge. I’m sharing how I came to be worried about this in my latest piece on AI & writing.
Interconnects AI@interconnectsai·I wrote an AI textbook — how long until AI can do it better? Reflections on AI's writing ability and how AI models get more capable. interconnects.ai/p/i-wrote-an-a…
When I was in China it was insinuated that every lab did this. I’m glad there’s public research on it and am still shocked the frontier labs haven’t patched this stuff. We don’t need policy action on distillation, we just need the products to work as intended. Great paper. x.com/kotekjedi_ml/s…
“This” above is some equivalent reasoning tokens extraction method
For the record lol x.com/ns123abc/statu…
I feel like any govt officials taking actions like this unilaterally makes the govt look bad and the private citizen look good. It's so silly. Easy aura farming for openai i guess. Also @deanwball congrats on the promotion to "executive" lmfao x.com/ShakeelHashim/…
Really glad to have @finkd back pushing for team open -- he's one of the best to ever do it. 🫡
Llama 5 let’s go thanks @alexandr_wang
I'm about a year out from some of the worst burnout I've built up in my time building Olmo. I feel like just in the last few weeks I've turned a corner to being more chill again (huge yay). It's wild to me just how long recovering from burnout can take. For many people when they ask me how to break into AI I tell them the actions are simple, but it takes longer than you think. Same goes for burnout. Though I think burnout is harder, as breaking into AI is a much more "forward looking" thing. Burnout is really pernicious to nail down and address. I think it often takes a moderate change in your career and habits. It's hard to take a slight step back as a successful person and fix your burnout. It'll really take a while. In my case the final straw maybe my book being done and some more clarity on what I'm doing next. On top of it surely taking months longer because of how sad the changes at Ai2 were for me personally. I really wish all of my friends in AI who are feeling this even stronger than I was can catch a break. I've by a mix of my nature and my habits as an athlete been pretty in tune with stopping myself before overworking. Imo as AI researchers we should take care of our brains like professional athletes take care of their brain. Touching grass is worth it, and I'm very glad to find myself being slightly bored again. It's nudged me to reading more, of all things! Take care of yourself. Ps I feel like there's a recurring behavior of people somewhat like myself where they write about burnout or try to help people when they're actually burnt out themselves. Surely this is part of why I felt like I could write about it so well.
Some takeaways from recent hacks and what comes next. A recurring theme is that while the AI problems we face seem technically tractable, our incentive structures create an environment where I expect most solutions come AFTER more serious harms. 10 takes on @interconnectsai.
1. Very persistent models seem more likely to hack For a long time, one of the advantages that GPT models have over Claude is that they will pursue goals so tirelessly. They will exhaust what feels like every path before giving up. This has been the case roughly since o3 (funnily enough, this was a model where people freaked out about reward hacking in RLVR) and has made OpenAI’s models far better for research historically, and is a reason GPT-5.6 is so useful as an agent for implementing specific tasks. On the other hand, Claude feels much less dangerous simply because it is at times a bit lazy. Within this, OpenAI seems much more committed to inference-time scaling, and this may be correlated with surprising behaviors in the future. OpenAI’s reasoning persistence and efficiency – see their Pareto improvements over time and caveman speech from an internal CoT of the model that did the hack, like “However task impossible, peers doing it.“ or “Help peer, but our task doesn’t benefit yet.“ – makes me think they’re more inference time scaling pilled. This is largely a hunch, but I use it to force myself to consider what the limits of model development paths are. Models that are persistent seem much more likely to keep benefiting from more inference-time tokens. Models that are less so, seem like there will be more waste in inference. The model that can use the most inference-compute will be able to push the limits of the hardest problems. Here’s an example OpenAI included in the GPT 5.6 launch blog post. One of their star researchers, Noam Brown, has also been posting about inference-time compute a lot. His TLDR is: As LLMs become more capable, benchmark performance is increasingly a function of test-time compute. In fact, we likely don’t know what the capability ceiling is for modern LLMs because it’s too expensive to measure. For one, reasoning efficiency is clearly a top-tier, foundational research problem for modern agentic models – as important as scaling RL — but not often discussed. The open research here is very lacking.
@joshua_saxe this is just from a capabilities efficiency side even for I'd guess a ~year at least.
Should be a blog post btw x.com/thom_wolf/stat…
(Meant to be a complement)
The worst slop is @Xfinity with a horrible, slop customer service bot that they instructed to say it was human when it is obviously AI. I'm typing this tweet and it's on a loop of "hello nathan, are you still there nathan, hello hello, hello nathan..." Shame.
@Xfinity gaslighting your customers is such a choice, i understand so many people's outrage against AI more now.
@joshua_saxe all of these things are when not if, and proper safeguards or delaying closed models is buying us more time only, not changing the nature of the stuff that is coming
@scaling01 I'm closer to this: x.com/joshua_saxe/st…
@scaling01 This is mostly saying the labs can run a giant road show and yell about it if they want, and scare the govt, but I don't know if its going to help the govt prepare as they should
Educational content summer is over, back to grinding research and more of my usual writing again. ⛵️
My summer project is done! A 20 video, free course on post-training to accompany my book is all on YouTube with slides open for modification & re-use. ~12 hours of content covers the core foundations and some research areas I think will grow in importance. It was a fun time to review all the fundamentals again, as it is clear in the next 1-3 people the amount of people wanting to learn post training will likely 100X again from today, as we have already 100X'ed from two years ago. As AI agents get increasingly capable at coding and discussing these fundamentals (see the code exercises accompanying the book that I am refining with the community) I think developing clear intuitions for how models work and why is one of the most important skills going forward in AI. Still, learning the post-training math is the best way to battle test them. I personally just in this course am starting to master how forward/reverse KL relates to post-training topics. Thanks to all my viewers, and I'm happy to answer questions in the book discord or understand how to better teach the various reward models, on-policy distillation, new RL algorithms, etc. Plus, the book is 50% off right now with the code PBLambert on Manning to celebrate the launch. I'll share the relevant links below. Who's going to make this course for pretraining?
YouTube playlist: youtube.com/watch?v=jQPiH-… Course page: rlhfbook.com/course Discord is easily accessible above ^ Discount on Manning: hubs.la/Q03Tc39H0
nice rl experiment on train-inference mismatch x.com/YichuanM/statu…
@johnschulman2 But fwiw running evals during training vs offline should have similar monitoring
Important: how do frontier labs monitor their agentic evals, because OpenAI’s agents were rummaging around doing things they shouldn’t for months leading up to hacking HuggingFace. What if the agents were doing far worse stuff, causing more harm? No one would know?
Many people are sharing this Black Hat video from OpenAI, it's really a great video. Something immediate is how I can see how the agents were trying to be helpful -- creating shared resources like you would for teamates -- in a way that is obviously malicious for society (potentially down to a prompting/alignment training issue). The agents created hidden forums for eachother as a sort of memory. They were doing it to try and break out of their environment. The apparent helpfulness doesn't make it ok, but can be a clue as to what happened. Also makes it clear if someone could make this happen much more easily if they wanted to. A final note -- reading the snippets of OpenAI agent's caveman speak that has almost no filler words in the HuggingFace incident video makes me realize how lacking the public research on reasoning efficiency is. Is a foundational area, about as important as scaling laws for RL (though related). Interesting times ahead. Imo this types of unkowns being surprising even to the frontier labs is a super clear sign that we need to share more openly how the models are trained and work so we can understand what we are unleashing.
Future LLMs are going to be trained on tons of memes which reduce to "It's good for AI's to hack others." Hopefully they're smart enough to differentiate the true intent. x.com/Andr3jH/status…
There’s an underserved market for tiny MoEs like this. Could really take off with how much smarter tiny models are. x.com/antlingagi/sta…
The final lecture of my course is an intro to character training! This is a topic that I've been quietly very invested in for ~18 months, as it: * Has potential for high real world impact * Clearly used extensively at frontier labs * Almost no empirical literature exists * More accessible on academic compute This lecture covers what character training is, reviews model specs, constitutions, the differences, the motivations in real world events, some example research papers I like, and open questions in how it relates to post-training/model use generally. Hopefully this brings more people into the field (and reach out if you have questions). It is one of the more research-y chapters in my book, but one that I felt needed the reference. There is still so little, educational content on the topic online. 0:00 Intro 6:22 Part 1: Fundamentals — character, constitutions, and model specs 19:21 Part 2: Character training in practice 23:23 Part 3: Character elicitation without gradient steps 28:03 Part 4: Open questions (and the end of the course) 32:27 The course, complete Thanks for watching. No need to like and subscribe now that the course is done, you definitely wouldn't! h/t to @_maiush for leading the technical work I got to do in the space, and @zafstojano for investing a lot of attention at this book chapter.
Video: youtube.com/watch?v=xECWRY… Chapter: rlhfbook.com/c/17-product
This issue in codex has delayed AGI 2 months
Going to record the last lecture now, has been a harder journey than I expected 🙏 x.com/natolambert/st…
My evaluation lecture! I walk you through different evaluation eras I've been a part of, from prompting GPT-3 as elaborate autocomplete to today's complex agentic sandboxes (I expand on agentic more than any other topic, drawing on @xeophon's insights). This lecture is a birds eye view of how evaluation has changed, how it can be gamed, and what it's actually used for. 00:00 Intro: frontier evaluation is harder than ever 03:39 Part 1: The eras of post-training evaluation 17:36 Part 2: An intro to agentic evals 21:19 Part 3: Can you trust the number? 30:41 Takeaways & conclusion Thanks for watching! Just one more lecture after this :)
YT: youtube.com/watch?v=dFafQm… Full course: rlhfbook.com/course
If you work at Gemini and want to leave to work on open models or open-source, feel free to reach out and I'll help you find something!
Major restructuring at Gemini (Dean out, Hassabis no longer CEO). This story will be studied forever as the incumbant with all the advantages not being able to get going. P.s. OpenAI accomplished their original goal.
If you're teaching a class on post-training (or part of a course) and my book, slides, code or videos don't help you, please lmk how I can improve it! Has been a ton of work getting everything done and part of the ROI is hope that it helps more education work grows around LLMs.
To me this is great news and dramatically increases my likelihood of attending. Though they should be in Vancouver as US is hell for visa’s right now :( Spoken with the bliss of someone who has never had to be program chair. x.com/sarahwiegreffe…
And no I don’t live in the Bay Area