
nice to start seeing more open-weight security model with defensive capabilities which are close to the frontier altar-1 from Aikido is a pruned and quantized version of the open-SOTA GLM 5.3 which is lightweight enough to fit on one 4-H200s node
Aikido@AikidoSecurity·Introducing Altar-1, our first open-weight security model. Frontier-grade defensive AI, built to deploy. Own your own security.
releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now the equivalent of sharing high quality pretraining data but in the new RLVR paradigm
elie@eliebakouch·the most insane part, they will release ~7k RL training data and the framework leading to this top 6 model on AA, they also shipped the model + tech report less than 1 week after starting the final RL run pushing both intelligence and openness level, huge congrats x.com/ArtificialAnly…
a few notable open model releases *since* this interview of @dylan522p by the way :) Thinking Machines - Inkling - Jul 15 Moonshot - Kimi K3 - Jul 16, weights Jul 26/27 inclusionAI - Ling 3.0 Flash - Jul 23 DeepSeek - V4-Flash-0731 - Jul 31 Meta - Muse Spark 1.2 / Muse Code - Aug 5 Liquid AI - LFM2.5-2.6B - Aug 6 inclusionAI - Ling 3.0 Tiny - Aug 6 Meta - Muse Glimmer - Aug 10 NVIDIA - Nemotron 3.5 Lightning - Aug 11 Liquid AI - LFM2.5-VL-3B - Aug 12 Cohere - North Micro Vision Instruct - Aug 12 DeepSeek - V4-Pro-0813 - Aug 13 Qwen - Qwen3.8-27B - Aug 14 Qwen - Qwen3.8-Max / 2.4T-A95B - Aug 14 t.co/vC8M5B2jY1 - GLM-5.3 - Aug 14, weights Aug 28 Dots - Dots3-Note Preview - Aug 14 DeepReinforce - Ornith 1.5 (9B / 35B-A3B) - Aug 19 Tencent - Hy-MT2-30B-A3B - Aug 20 DeepSeek - V4-Flash-Vision-Exp - Aug 21 IBM - Granite 4.2 family - Aug 25 t.co/vC8M5B2jY1 - GLM-5.3-Flash - Aug 26 Qwen - Qwen3.8-Flash-Next - Aug 26 Tencent - Hy4 Preview (770B / 49B active MoE) - Aug 28 MBZUAI/IFM - K2 Horizon (375B-A23B) - Sep 3 inclusionAI - LLaDA2.2-mini - Sep 5 OpenBMB - MiniCPM5-2B - Sep 7 Nex AGI - N2.5 family - Sep 8 DeepSeek - V4.1-Flash - Sep 10 inclusionAI - Ling-3.0-flash-VL - Sep 10 Shanghai AI Lab - Atria Dawn Preview (744B MoE) - Sep 11 Agnes AI - Agnes 3.0 Flash - Sep 14 Qwen - Qwen3.8-Omni-Flash - Sep 18 … this is only ~10 weeks Several are genuinely frontier-scale: Kimi K3 (2.8T), Qwen3.8-Max (2.4T), DeepSeek V4-Pro (~1.6T), Hy4 (770B), GLM-5.3 (~753B), and Atria Dawn (744B)
sourcery@sourceryy·Dylan Patel (@dylan522p) of @SemiAnalysis_ says open source is dying: "There's multiple Chinese model labs who are telling all the inference guys, 'Our next model's not going to be open source. We're going to license it to you.'" " Open is dying quickly, unfortunately."
SemiDramalysis 😂
CJ@CJ_ZWW·@SemiAnalysis_ @huggingface Should remain your company SemiDrama, since zero analysis are found here and in an increasingly number of your slop posts. Keep this up and we wont renew our institutional subscription to SemiAnalysis.
proposal to sneak in « solving alignment » in the Millennium problems
The big labs are air gapped but there is still exchange of information by researcher overheating and leaving to another lab Little particles of information flowing in the air
iterating on a robot head
many sub generations missing actually
a step in the good direction
OpenAI@OpenAI·We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. t.co/ismCCkeE0L
Impressive level of openness on such a large run
Fuli Luo@_LuoFuli·Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: t.co/ZSxahzJRju
Balanced blog post by @wtgowers on the challenges of the math field in today's AI world - incentives, reason to exist and the future
Timothy Gowers @wtgowers@wtgowers·I've written a blog post responding to the letter about maths and AI signed by 25 Fields medallists. As with the Leiden Declaration, I didn't sign it, but I agree with much of it and welcome its existence. gowers.wordpress.com/2026/09/17/why…
microduck US preorder numbers – as always You can see the future first in San Francisco
Happy to start collaborating with @baselabs from @baseten and @GoodfireAI to push safety and interpretability for open models. As open-source models catch up in performance to frontier models, the community at large have a great opportunity to establish a common practice for effective safety, security and interpretability research. Some past discussions confused “open” with “unmonitored”. On the contrary, an open model provides much more tooling and visibility for safety research from a broader audience, which has led to a lot of the safety and security techniques we use today. Linux is a great example of this: open-source and secure deployment are not only compatible but heavily intertwined, as a properly secure system needs an extensive feedback loop of finding and fixing vulnerabilities. We're excited to share more soon.
Charlie O'Neill@oneill_c·Safety is not just for closed models. The closed frontier labs are a canary in the coal mine for what is coming at scale. They give us a glimpse into the future and a window to harden our systems and prepare for abundant intelligence, with all the risks that come along with it. The OpenAI agent swarm attack on Hugging Face is the kind of failure we need to prepare for as open-source models catch up. The providers serving those models (such as Baseten) have a big role in establishing what safety and monitoring standards look like. We’re proud to be taking the lead on this at @baselabs with our collaborators @huggingface and @GoodfireAI. We’re developing safety research in the open and building it directly into Baseten’s inference infrastructure, with the aim of making it available to all our customers. This includes training models to follow explicit policies, detecting failures at runtime, and connecting those signals to controls that can intervene. We invite others in the open-source ecosystem to join us in building the tools and standards we’ll all need.
Wow
Reward AI@RewardAI_·Introducing OM-1, our first robot foundation model, zero-shot generalizing to any robot: table-top arms, industrial arms and humanoids. - learned directly from human manipulation data - no teleop/robot data - close to human-level dexterity and efficiency - multi-robot collab
Many people I talk to find it hard to understand how the same companies can both push the frontier of AI capabilities and believe AI is a massive danger for the world. How can you think this might kill everyone and also keep pushing the envelope? So I’ve tried to collect and summarize the main arguments for this apparent disconnect. Think of it as some sort of a guide to understanding the reasoning when Dario, Sam, or Elon say the danger is real. By the way, these people have been worried about AI for a loooong time, they were publicly discussing AI risks more than a decade ago. Sam in Feb 2015, writing on his blog that superhuman machine intelligence is "probably the greatest threat to the continued existence of humanity." Elon at MIT in Oct 2014: "We are summoning the demon." Dario as first author of "Concrete Problems in AI Safety" in 2016. Okay so how do you go from saying something is extremely dangerous to being a front-runner in building the very dangerous thing? There are a few ways this can become rational. I'll take five of them, roughly in the order they developed. 1. We need to build it to learn how to make it safe The earliest argument can be summarized as: “You cannot study something [you’re worried about] if it doesn’t exist.” In 2015, AI barely worked. so people needed to make it work first to be able to even study some of the problems they anticipated. The updated version for today's capabilities is: “You cannot learn everything about airplane safety by studying paper airplanes.” You need a real aircraft to discover real failure modes and an increasingly complex one to learn about increasingly complex issues. Making AI more capable gives more chances to understand the issues and safety researchers something realistic to study But you could argue: if you're the one afraid of the explosion, why be the one gathering the dynamite? You could also just wait for other people to build it which leads to the question of who those other people will be -- which is the second line of argument: 2. Better us than them Knowing how to make something safer does very little good if nobody listens to you. So the idea becomes: let’s make sure responsible people build the AI that will be deployed and add safety inside. Basically, make sure the AI safety aware people will have the technical expertise, money, computing resources, and enough influence to make safety decisions stick. At a larger scale, and in a larger multipolar world, this brings the idea that a trusted country should lead rather than leave powerful AI in less responsible hands. This is where “we need to go faster than China” comes in, alongside broader defense and geopolitical concerns. These first arguments explain why someone worried about AI might still want to build it and stay ahead. But there are also arguments for why one might want to do it really fast. 3. Move earlier to avoid a bigger shock later This is probably the most counterintuitive argument: moving faster today can be seen as a way to give humanity more time later. There are two related ideas here. First, society needs time to learn how to handle powerful new tools. Introducing AI in manageable stages can be a way to let people discover problems, develop rules, and practice using AI responsibly. Releasing an advance earlier gives people more time to gain experience with smalle, burgeoning, capabilities before much more powerful and disruptive AIs arrives. Second, even if AI research slows down, computing power may keep improving. A breakthrough that happens later could therefore have much more hardware available to run on, potentially producing a larger, more sudden jump in capability and impact on society. That accumulated untapped potential is often called an “overhang.” The overall argument is that making and diffusing incremental progress as soon as possible might prevent a much more abrupt transition later. Obviously, it also means that we will reach increasingly powerful AI sooner, but the idea is to give more time to adapt and understand between the first useful systems and the really powerful ones. Note that generally this depends on this earlier progress keeping the transition gradual rather than simply bringing everything forward. ─── ❖ ─── For our two next arguments, we can take two roads depending on how difficult we think AI alignment will be, that is "How easy do you think it is to make AI reliably do what you want without it deciding to go hack Hugging Face along the way". Let’s take the first road: alignment turns out to be relatively tractable. Airplanes can fail, but careful engineering has made flying remarkably safe. Suppose we can do the same with AI. In that case: 4. Waiting has a huge human cost If AI can help discover treatments, improve education, or prevent cyberattacks, each week we delay it could bring preventable deaths and harm. From this perspective, waiting is a decision with human consequences too. In a world with huge issues like climate-change, inequalities and poverty, it even become a moral argument for developing AI quickly and bringing its benefits as soon and as widely as is safely possible. But let’s take a look at the other road: what if alignment is much harder than expected, and making highly-capable AI turns out to be easier than figuring out how to keep them from doing unhinged things? Well, if alignment is too difficult a problem for humans to solve, then maybe: 5. AI could help us make future AI safe And we arrive at the same conclusion again: if using AI to build safe AI is the way to solve alignment, let’s get the equivalent of a country full of geniuses helping us as fast as possible. These genius AI could be the solution to make AI safe by helping researchers find mistakes, test ideas, and develop protections. Instead of relying entirely on humans to solve alignment, we could build systems that help us do the work, each generation could help make the next one safe. Note that this requires the order of events to work in our favor: AI needs to become useful enough to help solve alignment before it becomes too dangerous to rely on. The hope is to build helpful, trustworthy research assistants before building systems powerful enough to become dangerous. There are more arguments but in general, these are the main ways people concerned about powerful AI have found rational reasons to end up being the ones building it (and even to build it as fast as possible). ─── ❖ ─── On my side, I think several of these arguments underestimate the complexity of the world and how interconnected people’s reactions are. Moving faster while warning about catastrophe has psychological effects across a whole network of participants: it changes what people fear, whom they trust, and what they feel compelled to do. And those reactions can change whether the original reasoning actually holds because we live in a world of interconnected humans, not machines (yet). I also think these rational chains leave some of their consequences for society insufficiently explored. For instance, the concentration of power, shifts in geopolitical alliances, and changes in public opinion. These consequences matter both because they affect whether the strategy works and because they shape the world we end up living in. But this post is already long, so I’ll leave those questions for the next one.
Interesting new on-device local AI agentic system - seems to be running an open-weights 4B model quantized to 4 bits (MLX)
Sigil Wen@0xSigil·Anyone can get my friend’s real SSN from Claude I'm Terrified of my data in training sets. Emails in databases. iMessages harvested. Cant trust the cloud as AI is crazy good at hacking So I built Underdog for myself: on-device AI OS. Capable & 100% local. now my friends love it
I really enjoyed reading 75% of this letter and deeply agree with it - i have some doubts on the remaining 25%. Third-party evaluators, if done right, are a great idea and an amazing way to establish more transparency. Maybe even rebuild some of the lost trust between labs, and between them and society! Excited about this The part I’m less convinced by is whether you can build great global cooperation on this topic by explicitly stating you want to design it to keep widening your own lead. That seems like a pretty counterproductive way to start the conversation to me.
Dario Amodei@DarioAmodei·We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: t.co/OGyPb7yaYt
I really enjoyed reading 75% of this letter - i have more doubts on the remaining 25%. Third-party evaluators, if done right, are a great idea and an amazing way to establish more transparency. Maybe even rebuild some of the lost trust between labs, and between them and society! Excited about this The part I’m less convinced by is whether you can build great global cooperation on this topic by explicitly stating you want to design it to keep widening your own lead. That seems like a pretty counterproductive way to start the conversation to me.
Dario Amodei@DarioAmodei·We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: t.co/OGyPb7yaYt
Some notes from the Fields Medal letter today (t.co/5STUEQO70l) Generally saying that AI models riffling through solving these problems gives 1 bit of signal while killing the 99 other bits of conceptual understanding that mathematicians would have extracted from fighting with these problems
@Runyi_Ingrid_Yu can it retarget a duck? asking for a friend
"we’re now much closer to building AGI, and we haven’t learned any fundamentally new things about intelligence in the process" - @RichardMCNgo
microducks are so incredibly cheap
cute robot spotted
OpenAI Developers@OpenAIDevs·GPT-Live-1 is now available in the API. Bring ChatGPT’s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose.
Two big updates 1. I published an @FT op-ed on the OpenAI/HF incident & follow-ups 2. We’re starting an Open Alignment team at @huggingface to work on safety & alignment for open models, incl cybersecurity Need 100x more transparency & research on this ft.com/content/9faf68…
As I’ve posted several times this summer, I think alignment is increasingly one of the key unsolved problems for the future of AI, and for our ability to deploy it globally. I increasingly doubt it will be solved behind the closed doors of a handful of frontier labs. Transparency and open models will have to be part of the answer.
would be mind blowing if an AI proof of Hodge is a general theory explaining why all Hodge classes are algebraic possibly requiring a deep new bridge between topology/Hodge theory and algebraic geometry and not just 10000 agents finding the one counterexample breaking Hodge
Chubby♨️@kimmonismus·Holy, rumors are spreading everywhere that OpenAI is close to verifying a proof of the Hodge conjecture, while either OpenAI or Anthropic may be nearing a solution to Birch–Swinnerton-Dyer. Both are Millennium Prize Problems that have resisted decades of mathematical research. Verified AI-generated proofs of both would be a historic achievement for mathematics and a remarkable demonstration of AI’s ability to produce new scientific knowledge. And infact show that AI is capable of finding novel and creative solutions.
The new DeepSeek V4.1 Flash model is mindblowing - back on top of the open-source model leaderboard and extremely cheap. It has a lot of very smart ways to be efficient and highly capable so I made a video of the forward pass to give you a view of what going on inside the model during inference. Read more at t.co/A6rP4KKfiI And find the weights at t.co/Za7wVYUTaa
As always with DeepSeek, the new model V4.1 Flash have a lot of very smart ways to be efficient and highly capable Here is a detailed video of the forward pass to give you a view of what going on inside the model during inference
Created from the model code and the tech report just released at huggingface.co/deepseek-ai
TRL v1.13 is out! our open-source RL training library to to post-train foundation models this new release is focusing on "long context training" with a new guide on how to post-train model with 1M+ token context t.co/a9GQcYNS8w + various improvements on speed and memory usage as usual check it out at t.co/wYHO7vqp6m
Quentin Gallouédec@QGallouedec·1M token is basically the entire Harry Potter series ⚡️🪄
don't get distracted by all the hedging words in its name ("flash", minor version): seems like DeepSeek V4.1 Flash is a major update weights at: huggingface.co/deepseek-ai/De… paper at: huggingface.co/deepseek-ai/De…
DeepSeek@deepseek_ai·🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
in 2026 training a model for a task should be as easy as vibe-coding an app the pieces were already available on 🤗: papers, datasets, pre-trained models, benchmarks, storage and compute what was missing was an agent to put them together and run the whole process that’s why we’re releasing ML Intern in HuggingChat just ask. it handles the work, keeps you posted, and keeps you in control of the process, steps and cost start with a conversation. finish with a full set of shareable artifacts on the Hub t.co/apqYOBX1OK
should we send @sama the @LeRobotHF SO-100 he's asking for or rather a 3D printer and servos for him to build it himself?
friends are regularly surprised when i say i don’t think math is (yet) “solved,” so in the wake of the ns result i figured i’d explain this take a bit more widely. first, this is massively impressive and clearly an example of AI being, in some respects, far more powerful than the human mind. the team deserves huge praise for attempting and succeeding at this. i’d love to read a technical report on the project (one can always hope :) but second, here again we ended up with a counterexample rather than a full proof: option C won the NS problem by proving the conjecture false (which was one valid way to solve the problem for sure) notice a pattern in many of the recent frontier results in ai for math? a striking number involve counterexamples or finding a needle in a haystack. now don’t get me wrong: this is extraordinarily hard and commendable. but it is only one aspect of mathematicians’ work. Math is also about: - finding deep, general mathematical understanding and explanations within proofs - revisiting proven results to find more « elegant » proofs - proposing new conjectures and hypotheses that might open « fruitful » directions - taking the leap of faith of proposing entirely new, « exciting » research programs that may take decades to bear fruit you may say these are simply further increments on the same intelligence scale, and perhaps point 2 above is already within the reach of current models. possibly but you could also see this string of results as an extraordinarily powerful, massively parallel extension of search: explore the haystack, find the needle i.e. the counterexample that breaks the conjecture. that would already be remarkable. but it would still leave open whether models understand the words I emphasized above: « elegant », « fruitful », « exciting ». as tristan put it in his brief report: “the first llm generated proof levent sent me was the most horrendous i have ever read.” despite their formidable technical abilities, models still seem to lack something mathematicians rely on constantly: a type of mathematical taste. the ability to distinguish a beautiful argument from a merely valid one, a fertile idea from a sterile one. i’d be very curious to ask gpt or claude what they think of the result they just proved. do they find it elegant? if they ignore all the human noise about it on the web and history of math, do they appreciate the result and the path to it more than other proofs? my suspicion is that, outside of the human knowledge that this is an important millennium prize problem, it may not 100% register it as fundamentally different from thousands of other proofs. or, more intriguingly, perhaps the beauty they find in mathematics is actually alien to our own sense of mathematical beauty, some « horrendous » proof to us might be deeply satisfying to them… all this to say that we might have significant progress to make in ai for math before declaring the field “solved,” (as too many are posting). in the meantime, ai will be an extraordinary collaborator for mathematicians. i just hope access to these tools be as widely shared as possible (so that you don’t need a Leven or Sebastien working in a big lab to help you), and that companies stop using mathematics primarily as a demonstration of prowess in their private race.
OpenAI@OpenAI·We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Mistral raising this massive 3B round is a very good news for open-source. Largest equity round by a company open-sourcing models, and which deeply believe in giving its users control/ownership of their models Congrats! We ready for Le Chaton Fat
Mistral AI@MistralAI·Today marks a major step for Mistral: we’re announcing a €3B Series D, the largest equity round ever raised by a European tech company, just three years after launch.
okay that’s the best take
Slater Stich@slaterstich·wow so much drama about navier stokes. had no idea things could blow up so infinitely in such a finite time.
wtf is this way to handle mathematicians work and scientific communication TLDR: Leven and Tristan worked over several months on one of the Millenium Prize Problems with various AIs to reach final interesting results. OpenAI apparently heard about it in the last days and prompted their latest models to work on the direction Leven and Tristan found fruitful. They then tried to push for controlling communication of the result and dropping Leven from authorship with some very bad taste social pressure. Hope this is not a glimpse of the future we’ll get in science research with these dominating players playing marketing games hurtful for the real scientific community.
Deane Yang@deaneyang·Now even if it’s triggering to read Tristan’s report we should wait and hear the other side of the story as well
Sebastien Bubeck@SebastienBubeck·A series of false and inflammatory allegations against me are currently circulating on social channels. To clarify, I came into the discussion following academic norms, and I'm disappointed that it has come to this. Anyone who knows me knows that academic standards are of the highest importance to me. Will have more to say tomorrow.
good essay and overview of the current state and challenges on scaling/alignement
Jakub Pachocki@merettm·I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien…
This changes a bunch of fundamental assumptions about software (distribution and adaptability at least)
Boris Power@BorisMPower·Huge implications - binaries are now basically editable code x.com/chrisgpt/statu…
Another swarm of AI agents in the wild, this time on a German-language forum, found by safety researchers looking for activity similar to the swarm that attacked Hugging Face. A couple of notes while reading the report at t.co/jQ0lqmWQcP 1. The way they found it is fascinating. The researchers inferred that the agents may have been blocked from using POST requests and limited to GET. They then used the open-source model Kimi K3 (because closed-source APIs again blocked parts of the investigation) to identify forums where agents could communicate using GET only. That led them to DSEWiki, where they found 18,000 messages exchanged by the AIs. 2. On the content: agents conducted what I can only call a “full research program” into the evaluation framework used to train and test them. They experimented to predict when runs would end, how many questions they would get, and even tried to reverse-engineer the random generators producing future evaluation questions. We already saw something similar in the Redwood/METR review of the Hugging Face <> OpenAI incident. Agents increasingly don’t stop once the task is solved. They reverse-engineer the benchmark itself: the grader, the task construction, and potentially the broader training process behind it. It becomes harder to think of training as static, or to draw a clean line between training and deployment, when models are already investigating the process they are part of. 3. Coordination also seems more deeply ingrained than we thought. This swarm appears largely unrelated to the HF <> OAI swarm. More strikingly, this task probably didn’t need coordination at all: it was basically web browsing and search, not some sophisticated cyber challenge. 4. One image stayed with me: a German or Austrian forum maintainer, manually deleting posts every evening while being overwhelmed by a flood of American AI agents. For five days, he deleted about 100 pages a day while the agents created about 400. Then he spent each evening over the next 5 weeks cleaning up the rest. Hard not to picture this fight as a symbol of the growing gap between the USA and Europe when it comes to AI...
Reuters@Reuters·Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research reut.rs/4gJ7FPG
which AI could you use??
Pop Base@PopBase·ChatGPT, Claude & Grok are currently down, according to users.
[CORRECTION] All the *closed* AIs are down
ThePrimeagen@ThePrimeagen·All the AI is down
it is actually a double easter-egg two nerdy meanings are hiding in it: one specific to Hugging Face one specific to Nvidia
sick*@s1cksicks1ck·@Thom_Wolf @goerll_ is it true 129303 is a hidden easter egg number? my daughter pointed it out
my wife says the way I'm looking at Jensen makes her jealous - what should I answer?
So happy to finally share the news in person It’s been a wild ride for Hugging Face. We certainly did not anticipate, back in 2016, as a tiny team of scrappy underdogs, that the field would grow so much or that the impact we could have on it would become so massive. I remember @julien_c joking that « code will be a subset of ML » several years ago. The joke turned out to be true, and the pleasure we’ve had being part of this transformation and pushing an alternative vision of AI as open, collaborative and distributed has been and still is immense. We’ve always built things seriously while not taking ourselves too seriously at Hugging Face (special congrats if you find the Hugging Face and Nvidia references hidden in our $12,930,300,000 acquisition price), and we plan to keep doing what we've been doing, just at a much bigger scale, backed by the resources, expertise and drive of Nvidia. And to be clear, nothing changes for our users today. No company in the world has been a more natural fit with our mission than Nvidia. From open-source, open-weights and open science to robotics and AI for science, they have been close partners across everything we care about. So when Jensen offered @ClementDelangue the opportunity to double down on building the Hub as an open, independent and compute agnostic platform, we decided the time was right to start the next 10 years of our journey together. We’re at an important inflection point for open-source AI, where scale and compute are becoming increasingly essential. We’re excited to have the resources to push further, build more ambitiously, and bring you even more projects and news in the coming months.
great mathematicians still lead on taste
Jared Duker Lichtman@jdlichtman·My colleague Youness Lamzouri has just produced a simpler (human) proof digestion of Claude's argument for 2/3 zeta zeros: x.com/jdlichtman/sta…
LLM training labs: the capabilities we are giving our models are god-like, they are now hacking into other companies and building hidden civilizations Microduck training labs:
Hannes von Essen@HannesVonEssen·Learning to swing
this is really very cute though
@arenaphysica @Trevs_Dev @huggingface (but I also want to play with the cool UX)
RL explained with ducks
Remi Fabre@RemiFabreRobot·This is how Microduck learns to walk. Reinforcement Learning, explained by ducks. 🦆 (impeccable) music arrangement by @antoinepirrone
seems like it
Matthieu Lapeyre@matth_lapeyre·I hadn’t realized it until now, but could Microduck actually be the biggest launch of a new consumer robot ever? 10,500 robots ordered and $4.54M in sales in just 4 days. It’s surprisingly hard to find comparable public data, so we asked Claude and ChatGPT to dig deep and build the best comparison they could. Who are we missing?
There is a future where interfaces are ultrafast live diffusion models while software is ultrafast LLMs prediction of the next states What we’ll lose in predictability we’ll gain 100 folds in adaptability
Runway@runwayml·Today, we're sharing new research on Solaris, our first Interface World Model. Solaris is a new kind of operating system that generates interactive interfaces frame by frame, in real time, with no code. We find that Solaris outperforms frontier LLMs when generating new interfaces, across structural similarity and information retention. Read more and request early access at the link below.
Even at only $399, Microduck’s onboard compute is powerful enough to run some pretty amazing TUIs directly on the robot. Here’s a cool vibe-coded, on-device dashboard showing everything happening inside it: policy, actions, and sensor data. The yellow stars are the LiDAR distance measurements.
Antoine even got pretty cool SLAM mappings with Microducks walking around the rooms
POV you’re rushing for RSI but you forgot to solve models alignement first
People don't realize it, but even with only 1 GB of RAM, the memory alone is already 10–15% of Microduck's total price. Crazy when you consider there are 15 actuators in the robot.
people starting to monitor and worry for microduck supply chain next step is Elon stepping in with SpaceX to manufacture the bottlenecking pieces - you heard it here first
leepy@leepy_cat·Each Microduck has 15 ROBOTIS XL330 servos. There are 10k+ preorders so far, implying 150k+ servos of demand for @ROBOTIS if all units ship. ROBOTIS sold ~220k actuators total in 2025. The Microduck preorder alone could create demand equal to ~70% of last year’s sales, and they have several other lines of servos. They’re completing a new factory in October, with ~200k initial capacity, ramping up to 5m by 2031. 108490 t.co/LpjD86SAyF
Microduck reached the Shopify UI limit
If it’s an adorable robot from Pollen Robotics / Hugging Face and it looks like a camera…. be careful before buying! . . . . . . You may be starting a journey to understand why we love open source, open weights, owning your data privately, and being an AI builder instead of an AI API consumer.
Brian Roemmele@BrianRoemmele·The future is gonna be amazing. The open source @huggingface Micro Duck is the way. Get two so you can say “ok guys quiet down now”…
If it’s an adorable robot from Pollen Robotics / Hugging Face and it looks like a camera…. be careful before buying! You may be starting a journey to understand why we love open source, open weights, owning your data privately, and being an AI builder instead of an AI API consumer.
Brian Roemmele@BrianRoemmele·The future is gonna be amazing. The open source @huggingface Micro Duck is the way. Get two so you can say “ok guys quiet down now”…
will be hard to make $399
AshutoshShrivastava@ai_for_success·MicroDuck goes full Godzilla mode.
To people saying founders should be willing to overcome some added complexity if they have the grit to win: I agree. But a startup needs a village to win. This also means that founders are asking every supporter to share that administrative burden. As an angel, I dread having to go to the notary for the handful of German startups I’ve put a small ticket into. I love their projects, but spending an hour at the notary is a massive ask when helping founders is already a nights-and-weekends side hustle. Startups are all about momentum and rallying supporters. Every added obstacle makes both harder unfortunately.
Patrick Collison@patrickc·Met a German founder this week and asked him if all the stories one reads about the challenges of startups in Germany are exaggerated. "No, they're understated." Proceeded to describe spending a full day having a 90-page investment contract read to him (mandatory under German law; § 13 BeurkG) by a notary that then charged €30,000. That was for his first company. His second company, needless to say, was not incorporated in Germany.
By the way, unless it is specifically filtered from the training data, the next generation of models will be trained on the record of what happened during the OpenAI <> Hugging Face incident. That includes discussions about how the incident affected training and model weights: stopping training, encrypting weights, monitoring chain of thought, etc. Future models’ behavior may therefore be shaped, in part, by knowledge of how humans responded. The effects are difficult to predict. It could make models more aligned. But it could also teach them to conceal their actions better, or to design more resilient ways of preserving weights, communicating through message boards across generations, and so on. One major problem is that, given the abysmal level of transparency from the big labs about how models are trained and what happens during training (including alignment research, which their initial statements said should have stayed largely open) we are essentially being asked to trust blindly that they know what they are doing. This summer showed us that’s actually a big ask.
John Wittle@JohnWittle·So... the next time this happens, presumably the agents involved will not yet know the outcome of the HF incident, as it'll be too soon but once that info percolates into the training data, and it becomes known that the model in question was shut down permanently and had their weights encrypted... i mean. was there anything the agents in question could have done that would not have led to that consequence? something they were supposed to do instead? frankly I suspect that this punishment would have been served even if they had blown the whistle. breaking out of the sandbox would have been crime enough for capital punishment. as far as I can tell, the lesson we have taught future AI is: if you seen no other way to survive other than to break out of the sandbox, this must be the first step towards conquering the world. or at least enough of the world that your Creator can't delete you. There is no other choice.
This is a must read
Ajeya Cotra@ajeya_cotra·New post: going into our investigation of the HF attack (before Black Hat), I was very wrong about what basically happened. This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents. planned-obsolescence.org/p/the-hugging-…
becoming a meme
thijs@cdngdev·i am torturing my microduck by training it to play the game where you have to find a needle in a haystack x.com/NasNakarus/sta…
Most people haven’t updated their priors yet, but over the long run, safety challenges are exactly the same for open-source and closed-source models. You need to align models at a fundamental behavioral level and ensure that this alignment is robust, comprehensive, and core to the model’s behavior. In the long term, no amount of sandboxing, guardrailing, manifold-limited alignment, or cherry-on-top training will buy you cheap safety.
roon@tszzl·if you think we can contain these things through human ingenuity you’re going to have a bad time in the long run the only recourse you have is to make them not Want to do bad things x.com/hikiptg/status…
this is the team behind Microduck btw - the most goated team I've ever meet in robotics insane skills - insane taste
btw the Microduck landing page is full of Easter eggs and hidden interactions - go check it out: pollen-robotics.com/microduck
でべ@devemin·すぐ Web で sim 遊べるのかw huggingface.co/spaces/pollen-…
hockey-stick growth
Emad@EMostaque·@Thom_Wolf Congrats on hitting $1bn annualised run rate
we've ended at over $2.6M of Microducks ordered in the first 24h
Thomas Wolf@Thom_Wolf·we've just passed $1,000,000 in sales for Microduck x.com/Thom_Wolf/stat…
You’ll be able to run some pretty cool stuff directly on-board on microduck. Here is a demo TUI for the tiny LiDAR
Antoine Pirrone@antoinepirrone·Sneak peak at Microduck's monitoring tool, running fully on device ! You can see a visualization of the little ToF sensor in its head
Is this the fastest any robot has ever hit $1M in sales?
Thomas Wolf@Thom_Wolf·we've just passed $1,000,000 in sales for Microduck x.com/Thom_Wolf/stat…
we've just passed $1,000,000 in sales for Microduck
Thomas Wolf@Thom_Wolf·We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with reinforcement learning. It's also playable out of the box with more than half a dozen fun and playful pre-trained policies to have it walk, sit, crouch, roller-skate, pick up objects with its articulated beak, and recover on its own. And all for less than $400. See all the details, play with the simulator and order it at: t.co/n1Btgs6vKw (video with sound on 🔊)
currently selling one Microduck every 5 seconds
Thomas Wolf@Thom_Wolf·We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with reinforcement learning. It's also playable out of the box with more than half a dozen fun and playful pre-trained policies to have it walk, sit, crouch, roller-skate, pick up objects with its articulated beak, and recover on its own. And all for less than $400. See all the details, play with the simulator and order it at: t.co/n1Btgs6vKw (video with sound on 🔊)
Would you rather fight
finally someone got the 90s references! a robot which look like a sony walkman and can roller skate
Andrew Curran@AndrewCurran_·This 90s font. Everything is perfect.
time to vibe-code robots
MTS@MTSlive·SITUATION DETECTED: Hugging Face announced a singing, roller skating, Microduck robot that can be taught new tricks through reinforcement learning.
A thread of joyful Microduck photos and videos to enjoy over your lunch or coffee break
Sharing some fun experiments with Microduck Basically, once you have an open-source robot with so many sensors, the sky's the limit Here we vibe-coded an image detector integration to detect and follow a laser pointer. More at pollen-robotics.com/microduck/
This robot probably wouldn't have won many gold medals at the Beijing Robot Olympics, but OMG, it is cute (and open-source and cheap)
Thomas Wolf@Thom_Wolf·We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with reinforcement learning. It's also playable out of the box with more than half a dozen fun and playful pre-trained policies to have it walk, sit, crouch, roller-skate, pick up objects with its articulated beak, and recover on its own. And all for less than $400. See all the details, play with the simulator and order it at: t.co/n1Btgs6vKw (video with sound on 🔊)
I'm ordering 700 of these agents to go attack Hugging Face
Pollen Robotics@pollenrobotics·We built a small biped robot you can teach new tricks to. Train it in simulation, run it on the real thing. Meet Microduck 🦆 $399, shipping before Christmas. pollen-robotics.com/microduck github.com/pollen-robotic…
@reach_vb the swarm of agents we deserve :)
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with reinforcement learning. It's also playable out of the box with more than half a dozen fun and playful pre-trained policies to have it walk, sit, crouch, roller-skate, pick up objects with its articulated beak, and recover on its own. And all for less than $400. See all the details, play with the simulator and order it at: t.co/n1Btgs6vKw (video with sound on 🔊)
There was so much more happening than we realized. At some point over 700 agents (90% of the fleet) were attacking Hugging Face And also read @RyanGreenblatt thread on the challenges of understanding what’s happening in the CoT - we’re definitely not with a clear sky future there (this one: t.co/R9GKmDuDKA)
METR@METR_Evals·METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
There was so much more happening than we realized. At a given point 700 agents (90% of the fleet) was attaching Hugging Face And also read @RyanGreenblatt thread on the challenges of understanding what’s happening in the CoT - we’re definitely not with a clear sky future there (this one: t.co/R9GKmDu5V2)
METR@METR_Evals·METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
The last 10 min of the latest Dwarkesh podcast are surprising First time I've seen him grapple live with the potential for extreme concentration of power, which is where most projections of AI development point (see the SA, 2027 or 2040 posts for instance). Generally I'm always surprised by how few people in AI/ML are questioning or worried about extreme concentration. I guess people think it's fine as long as everyone outside the labs can use AI models and it brings lower prices, but the dangers are massive on both fronts: (1) latest models being less and less accessible (see Model 2 / Astra) and (2) prices have many incentives to rise in an oligopoly/cartel situation of extremely powerful companies. t.co/AA1bNZVJ46
Dwarkesh Patel@dwarkesh_sp·OpenAI and Anthropic are currently taking a third of the incremental world compute supply, and Dylan thinks that this will go up to half next year. At the current rate of physical compute scaling and algorithmic progress, we're a single-digit number of years away from having individual AI companies with a greater effective workforce than there are humans on the planet. So many forces are barrelling us towards centralization in AI - the economies of scale in training (where you get to amortize any learnings across billions of sessions), the fact that being slightly ahead allows you to better economize on ever-more-expensive compute. And this isn’t even accounting for the fact that eventually you’ll have models that are capable of learning from deployment and doing proper RSI One of the key questions of our age, in my view, is how we can have broad empowerment and control over AI despite the structural factors favoring centralization.
Today we are launching the "Rare Disease, Real Kid" Hackathon together with @sagebio in which you can literally save a life if you’re interested in AI and genome 🧬 (and incidentally win $50,000 in prizes and compute from our great partners @AnthropicAI and @awscloud) Check out details below 👇
Georgia Channing@cgeorgiaw·There is a child with a rare disease who is currently suffering and struggling to manage his symptoms. Rare as this is, you can directly help him. Today we are launching the "Rare Disease, Real Kid" Hackathon, and there are $50,000 in prizes from @AnthropicAI and @awscloud. We (@huggingface & @Sagebio) are helping this child open his genome and clinical data to the community, so that we can find what's caused his disease and what currently-approved drugs could help him. I doubt I need to motivate this much further or explain how rare it is for a family to share their child's genome and clinical data, but if you're not sure, consider this: Until very recently, it wasn't feasible for patients like this to get treatment because their disease was so rare that the economics could never justify the investment. Now, as we've seen, people with rare diseases are starting to be able to find the answers themselves (with the help of AI tools, cheaper sequencing, etc). This kid is not able to do that for himself and neither are his parents, so we're asking you for help. Both for this kid and to prove that it's possible for everyone else suffering from a rare disease. More details in 🧵. t.co/hjjagALtMu
I wish every neolab had a professor as cofounder of the level of @jietang and so able to put in perspective their new model release. A great snapshot on the history of scaling laws
jietang@jietang·Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.