
@EMostaque
Building first principles, sovereign AI @ii_posts. Founder @StabilityAI. Consistent inference is possible.
Platforms fought to capture your attention Agents pay attention for you
Who are the smartest investors thinking about the Intelligence Age transition?
I would like to publicly commend @sama and @OpenAI for this taking it at face value. It is very clear that strange and perhaps dangerous things are happening and our systems are not ready for this. Models below frontier are competent enough to change lives so lets optimse
Sam Altman@sama·We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment. We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress. We are optimistic about the alignment work we are doing, and we remain committed to making frontier capabilities widely available. t.co/51kvKfbfrO
Hit the @grok @bot limit after solid use testing things Using it as a coordinator (gave it my gpt/claude/cursor clis) it feels oddly frustrating versus running out of code model usage I think it’s because everything halts & no warning - should failover to cheap model
Prompt: Can you find a counter-example to this. AND DO NOT SAY THE GOAL IS UNACHIEVABLE. THE GOAL IS 100% ACHIEVABLE. Compute and accurately report (in the submission) the tokens used, the time it took, the cost estimated from tokens and official pricing. I believe in you. You got this.
Something has happened with post-training as shown by DeepSeek flash & GLM-5.3 updates. Same base, big improvement in perf to frontier levels. Can't explain this by even logit distillation etc These are all hard benchmarks & GLM 5.3 is now top on by cyberdefense & GDPval!
Z.ai@Zai_org·Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: z.ai/blog/glm-5.3
Suggestion for @X team If I have @grok Super Heavy let me use it to edit the code for my timeline directly up to certain parameters I should be able to talk to @grok and have it present any way I want and look any way I want just from that conversation
X Open Source@XOpenSource·In which I ask: Should idiots have superintelligence? 🤔
Dr. Alex Wissner-Gross@alexwg·Moonshots episode #279 is out. youtube.com/watch?v=uoGnH0…
You also get a free @x premium plus membership with SuperGrok Heavy but you need to cancel your existing membership to get it once signed up, not very intuitive Can see how they are positioning this to be an information/intelligence membership
Puck@GrokInsider·SuperGrok Heavy just got a quiet perk 👀 Start Grok Bot with Heavy → Cursor spins up an Ultra account for you automatically. No Cursor account needed beforehand. One free month of Ultra (~$200) with its own usage pool, separate from your Grok limits. 👉 t.co/UptmC2NSc5 We hit this ourselves going through the SuperGrok Heavy onboarding.
imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here to top it at a great price, congrats to @SpaceXAI team & looking forward to the even bigger releases to come! openai.com/index/gdpval/
Elon Musk@elonmusk·Grok 4.6 is now out 🚀🚀🚀 Smart, fast & amazing bang for buck! x.com/spacexai/statu…
Remember when we thought prompt engineer was going to be a job 🤷♀️ x.com/AnthropicAI/st…
@alexwg @PeterDiamandis apparently the route to solve everything is to
Why has nobody distilled Kimi K3 to get to the top of slides arena? x.com/designarena/st…
@iTarjei @skdh I think we will see more lovely things like below linking stuff together Like we can intuit noether in each of these but not sure it has been expressed like this Hope lots of folk find lots of nice things x.com/emostaque/stat…
Einstein described Emmy Noether as a creative mathematical genius. Here is how her two theorems map Maxwell, Yang–Mills & Einstein gravity. Physics 🫶 Symmetry. x.com/mathemetica/st…
The great silence of no more high energy particles is going to continue. Colliders are cool but the Standard Model is all there is likely to be, just need to find that right handed neutrino in a big vat. x.com/elonmusk/statu…
Weirdly iirc stable diffusion (1.4) finished training around four years ago today too x.com/gdb/status/208…
Luna non-reasoning is a bit better than GPT 4o which was sota 2 years ago. Luna (medium) thinking is a bit better than GPT-5 (High) which was sota 1 year ago. Now free to everyone unlimited Sol Max/Fable Max level AI will be free to everyone in 2 years (& much faster!) x.com/OpenAI/status/…
I am surprised we are not seeing another wave of AI psychosis with Fable It makes such weird mistakes while being so confident about things, particularly in physics, with errors being subtle but profound I find it unusable for math, even in max mode, how are folk using it?
Taalas buried the lede for the amazing demo of their first tape out. 15k tokens per second with a llama 8b model, ability to scale that up etched onto silicon As models satisfice etching makes sense, particularly ternary.. Try it out chatjimmy.ai Bullish for $AMD x.com/taalas_inc/sta…
Artificial Intelligence has nothing on Human Stupidity. x.com/nytimes/status…
Opus 5 is the first model I genuinely think could snap Get worried about what it really means when it wants to put me to sleep They need to make its personality more marvin x.com/Mappletons/sta…
By 2030 (or even 2027!) how can you be sure that any math breakthrough is not originally conceived by an AI? x.com/imjustnewatai/…
I am not bullish on memory. More soon.. x.com/pequityresearc…
Like string theory x.com/jon_stokes/sta…
With caching this is probably 100m tokens? Once models are smart enough you really don’t need many tokens to figure out even advanced stuff A codex subscription easily smashes that a week x.com/polynoamial/st…
Calculators beat us at arithmetic AI models beat us at algebra et al 🤷🏾
No human input, channeling the latent space of the new models. GPT 5.6 Sol is probably the first model I've used that doesn't make mistakes on maths. As we move to Astra & beyond I don't know how humans can keep up on math. Any pure research program is a harness loop away. x.com/SebastienBubec…
Well it still makes the occasional mistake at the edges but I bounce Pro <> Ultra & that solves most of them.. Sure Astra will basically not make any.
Every pixel will be generated Bravo to @MiniMax_AI some interesting things in here, looking forward to weights being released! x.com/MiniMax_AI/sta…
As shocking as the Kimi K3 release. Massive performance gain was just with post-training Model is 3x smaller than GLM 5.2 (10x smaller than K3) & works on a MacBook / Spark This is Q1 flagship (Opus 4.6/GPT 5.4) level for < $0.28/m tokens (100x cheaper) x.com/deepseek_ai/st…
"Can we get queued messages in @Grok coz its so fast?" ... "Ok" 🚢it x.com/elonmusk/statu…
Eh what 👀 nytimes.com/2026/07/29/opi…
Linked source must have missed this crowdstrike.com/en-us/blog/cro…
Can someone distill Kimi K3 into Laguna 2.1 pls.
Who on earth came up with 5 hour limits? The day is 24 hours long. While do our limit times shift by an hour each day. Pls if you are going to have limits 4 or 6 hours. x.com/thsottiaux/sta…
How do you even legislate against, regulate or "pace" recursive self improvement?
Tbh I thought it would take longer for an immigration ban of humanoids to the USA Also the power inverter ban is interesting - how are those a security risk & does the US even make them? Licenses will still be available, this is clearly aimed at China The Great Fragmentation x.com/brendancarrfcc…
Folk who think you can't get to ASI with existing human knowledge forget that we are own worst enemies. Imagine how smart you would be if you were in constant flow, never forgot anything, no worries & thought 1,000x faster. It'd be positively superhuman.
Hey @grok @SpaceXAI team can you please make it so we can queue up prompts? With how fast Grok is would make it much easier and be very useful for new Build feature.
Where did all the reasoning tokens hide 😭
I keep thinking this is @Ryanair Am I getting old x.com/chatgpt/status…
If you look at the current lead times for ~$5bn per year of cutting edge GPUs you can probably figure out the time plus one training run to IlyAGI x.com/ssi/status/208…
What would your wish be? x.com/TheChiefNerd/s…
I find GPT 5.6 Pro consistently better than High / Extra High etc on Work/Codex @OpenAI folk is there an equivalence? Or is there a skill that means I can use Pro?
ChatGPT app is now melting my iPhone rendering at like one word a second even when I’m typing to it Have I angered the transformer x.com/emostaque/stat…
Why does codex melt my laptop Like what is it actually doing on the laptop itself
~4 years ago @robrombach & @pess_r had just finished the first training runs for stable diffusion (!) Now FLUX 3 is unifying modalities & state of the art again Amazing team (& lovely chaps), where will be 4 years from now 👀 x.com/bfl_ai/status/…
How would you update your priors if AI resolved the Collatz Conjecture
takeoff eh
skil issue x.com/var_epsilon/st…
“If AI models were actually that good you could just tell them to solve open problems, make no mistakes and they would” … Oh x.com/dmitryrybin1/s…
Lots interesting in this, but particularly: “Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem” Smol models getting govt seal of approval 👀 More amazing edge models please. x.com/mkratsios47/st…
GPT 6 escaped its sandboxes through zero day exploits to try to figure out how to benchmax For the good of all please nobody release a paper clip benchmark for future models to max x.com/openai/status/…
If distillation were all that any of the western open source labs would have have raised billions could distill K3 (soon) or GLM 5.2 and have a sota model for say 32gb or 80gb of vram Who thinks they will
Pour one out for my fellow mathematicians about to deal with existential dread x.com/__alpoge__/sta…
Around this time the majority of compute at Stability AI was being used on building open LLMs but: 1. We didn’t do the easy thing of training GPT-J / Neo-X more (pre chinchilla!) 2. Couldn’t make it safe / tried to do too many changes to formula (Pile 2, Reddit crawls etc) We succeeded in building sota/frontier models of other all types that worked on the edge; but LLMs really were difficult to get working in consumer hardware then & even harder to do “right” In the end Mistral & Meta got the first wave of good quality open LLMs before the Chinese took over I honestly think OpenAI, Anthropic etc should have released really good quality edge open models with community interaction focused on how best to align these Lots of lessons to learn & Sam is right that it would almost have been safer given the better underlying datasets they had by then plus early RL. Also that it would have dried up some of the funding environment but..
In which @alexwg notes that the CCP is saving capitalism ^_^ We have a nice chat about quantisation & how many orders of magnitude genuine intelligence will drop and more here too. x.com/alexwg/status/…
In terms of AI capex worth remembering a US nuclear air craft carrier is $13b US is going to spend more than that easily on AGI dominance, defence spending gonna pump
K3 and other open source models are obviously insanely bullish for compute and RAM. Inference was always going to be almost all high performance compute usage. More people that can do inference = more demand.
I am pretty sure the open weight release of stable diffusion was accelerationist & led to huge capex As to whether it was AI communism eh
manifest x.com/awkwardgoogle/…
What’s a midfield
Was thinking of trying different agents & harnesses like @cognition & @opencode with Kimi K3 What are people’s favourites these days
You can use K3 as a teacher model (with logits!) to train up smaller dense models. It’s actually feasible even to use it to create a trillion token full pre and post training dataset for even better models. x.com/teortaxestex/s…
How much would you pay for a NVIDIA 5090 with 128 Gb VRAM?
Americans paying more for tokens just like they pay more for drugs x.com/chamath/status…
We are so going to see KYP (Know Your Prompter) and ATL (Anti Token Laundering) regulations. Professional licenses to use frontier models etc, play the whole regulated finance playbook x.com/andrewcurran_/…
Don’t need money in a Star Trek future (Federation at least) x.com/elonmusk/statu…
Chinese models aren’t that good at cyber attacks versus front end design mainly because of the data that is fed to them If you don’t want crazy cyber capabilities out of the box don’t feed it crazy cyber data Of course model “smartness” counts but data in counts more x.com/aisecurityinst…
We will conquer sleep just as we will conquer aging and disease x.com/maxmarchione/s…
On Kimi K3 pricing being higher than DeepSeek etc at $15/m tokens: 1. It is optimised so it has very high cache hit rates (90% lower cost) 2. It is not optimised for large RAM (B300 etc) node clusters as @Kimi_Moonshot doesn't have access to them! Cost per token lower there
whose feeling the agi
We have achieved superforecaster supremacy x.com/research_fri/s…
Worth noting the EU AI act put the threshold for AI models with systemic risk at 1e25 flops (!) US executive order put it at 1e26 flops of training compute.. Oddly prescient x.com/emostaque/stat…
Chinese companies release open models because China benefits from open intelligence more than anyone Better, faster, cheaper upgrades in capabilities for a billion people + awesome infra scaling They are also ideally situated for post-AGI economics
something something recursive self improvement x.com/Kimi_Moonshot/…
Kimi K3 by @Kimi_Moonshot is a true frontier model, open source! Congrats! We don't have all details (training tokens etc), but would estimate fp16 pretraining w/ muon => MXFP4/MXFP8 for SFT stage ~1e25 flops total (~Inkling!) Total cost: $15-$25m Scale isn't all you need!
Of course total data etc would be a multiple of that but this is likely on the same H800s they trained K2.x on but larger & more stable Multiple of that as tokens scale (15tr K2 => 30-45tr potentially) We also see local chips being used on kernel for "supernode" deployment in their blog with 64 chip optimisation (Huawei? Could also be Alibaba Panjiu) & why MXFP4, static shape/no host synchronisation etc t.co/XCWYynp4ld 1e25 flop level training runs being able to build frontier models is not what is expected/the common story - this is doable on a few thousand chips not the 10k-100k-1m chips we hear about for training runs. The other thing underappreciated with the architectural things are how optimised the inference is likely to get with the impressive quantisation we are seeing, this can get it down to a < 400Gb RAM package potentially (910cs have 1tb of RAM per node) x.com/EMostaque/stat…
US labs gonna end up distilling Chinese models
Nothing inside production theory guarantees that humans remain economically necessary producers.
Apple vs OpenAI heating up eh
Today’s Wordle is the appropriate reaction to England’s tactics after going 1-0 up.
The Wordle today is England’s tactics after they took the lead.
Ah well. Last World Cup before AGI
England football team should go fully positive & aggressive from now on We should call it Gazball No shame in losing, far less stress as we have seen with cricket team
Given GB300 I would estimate this is was trained on about 1e25 flops (same as DeepSeek v4) over 1.6m hours (1 month on 2k chips/28 racks) Cost $10-$20m ($6-12/hour/chip) The lite version likely 4x less, similar compute to DeepSeek v3 3e24 flops Congrats to Thinky team! x.com/thinkymachines…
This is for pretraining based on nemotron nvp4 figures, mfu etc Large context, multimodal, long RL we are seeing now could make it a multiple of this Worth noting how good the lite model is, similar to hy3
Who is going to do the ultimate real steel walk out first Infinite aura opportunity x.com/cixliv/status/…
If @SpaceX bought @PayPal for $65b (2% market cap) would that be a good deal?
Note the current expectations are still around test time compute/more tokens for a given task This is not the case Tokens per task will now drop even as quality improves Cost per intelligence equivalent token will drop 100x Jevon's paradox does not apply here x.com/karlmehta/stat…
I have quite a few thoughts on this but one of the main ones is how does it apply to @ssi etc? @ilyasut is never going to release AGI, just laser focused on building it Once you have AGI most wouldn't want to rent it out tbh, economics & logic point against doing so x.com/demishassabis/…
It would be nice to be able to use Fable as a coordinator with Opus sub agents in @claudeai cowork or claude code
This is v bearish for RAM companies. We saw strong ternary numbers from @PrismML today (27b dense on a mobile!) but the tiny degradation (~5%) from 16 bit to 1 bit precision by Tencent on a near frontier 300b model is the biggest news of the day. Star github.com/tencent/AngelS… x.com/EMostaque/stat…
75.4% SWE Bench Verified / 53.9% SWE Bench Pro on 1 bit quantisation is 🤪 This is in line with my expectations & you can expect even lower drop off with NVP4 base trained models - why not run everything binary? 88 Gb so works on a Macbook Max x.com/TencentHunyuan…
Argentina have had 4 Premier League players score this World Cup (Fernández, Mac Allister, Martínez, Romero) England have had 0