
@SemiAnalysis_
sourcery@sourceryy·Dylan Patel (@dylan522p) of @SemiAnalysis_ says open source is dying: "There's multiple Chinese model labs who are telling all the inference guys, 'Our next model's not going to be open source. We're going to license it to you.'" " Open is dying quickly, unfortunately."
The datacenter moratorium curve is rising. But where does the risk actually sit? (1/5)🧵
Our September 15 report: t.co/lkbw4X4GQ1 offers a useful breakdown: 🟠 400+ local actions and proposals tracked across all statuses. 🟠 300+ local moratoriums enacted. 🟠 20 GW inside live local moratorium boundaries. 🟠 1.5 GW estimated to face actual delays from those restrictions. That means roughly 7.6% of nominally exposed capacity is delayed. (2/5)
Huawei has extremely strong and passionate, yet kind and friendly, software and hardware engineers. Their open-source Ascend inference engines are making good progress!
Huawei has extremely strong & passionate yet kind & friendly software & hardware engineers. Their open source Ascend inference engines are making good progress and proud of the work they are doing!
Is EMIB really chipping away at CoWoS? CoWoS has owned advanced packaging for years. TSMC keeps adding capacity aggressively but it still isn't enough to satisfy customer demand. Intel stepped in with a compelling alternative with EMIB — and customers are responding. (1/3)🧵
ChipBook August data caught a +467% YoY jump in HBM exports from Korea to Malaysia, ~$3B in just two months. We attribute this primarily to Intel's advanced packaging capacity in Penang, with LinkedIn showing ~50 back-end job openings in Malaysia. Not an insignificant number. Exports to Taiwan slowed in return. Trade flows and job openings are starting to confirm customer shift. (2/3)
Computation and Data Movement for Inference Mapping MoE models onto inference hardware: structure, flow, and efficient serving newsletter.semianalysis.com/p/computation-…
Experiments have shown that offloading Engram to DRAM results in up to 50% better performance on H200, B200, B300, and even GB300 NVL72. Our technical team has collaborated with 10x engineers from @EmadBarsoumPi & @AnushElangovan to upstream support for @vllm_project ROCm Engram DRAM offloading too, and we are already seeing up to 50% better perf too in PR 57491. As such, in PRs 985, 1002, 1005, 1006 @NVIDIA, @AIatAMD and @SemiAnalysis_ are updating the default recommended recipes to use Engram offloading.
Neocloud is now an entire industry. The word was coined by a SemiAnalysis researcher. Some of the companies using it still have no idea. "Nobody believes me when I say that. I'll tell people externally, yeah, we're the ones who invented the word, and they're like, No, you didn't. I'm like, Yeah, great." "This guy didn't know us at all. He was like, Cool man, what do you guys do? I'm like, What the f**k? We invented your entire industry."
Watch Now: youtu.be/Qv2vMj0Uq3c?si…
We took the silicon out of the silicon. Quick turn from SemiAnalysis STEEL Teardown Lab: iPhone 18 Pro Max, A20 silicon on TSMC N2. more to follow...
Ever since NVIDIA acquired @HuggingFace, we have been looking into migrating some of our work off of HuggingFace and to alternative solutions like ModelScope. Even though NVIDIA's announcement claims they will continue allowing HuggingFace to be accelerator-agnostic, NVIDIA does not have a good track record of developing hardware-agnostic software. We love HuggingFace and hope we are wrong, but at the same time, we are also finding the UX of ModelScope to be great!
If the claim in this article that Dylan and Dwarkesh aren't related is true, then explain this:
CAT@cattttang·I sat down with @dylan522p and spoke with over a dozen people about his information machine. "I don't think anyone who has a traditional equity research background would even try to figure out where $3T in [the AI labs' combined] ARR comes from." Slack private messaging limits, finding CoWoS volume, and the industry they're looking at next - I wrote about SemiAnalysis here: t.co/0FVDIJNaGe
More than 300 US local governments have voted to pause datacenters in the past 18 months, and the pace picked up sharply this summer. Most of them are buying time. "A datacenter moratorium is a temporary legal pause, usually enacted by a local government. It halts approval, permitting, and construction of new datacenters. It doesn't affect existing datacenters." "These can last anywhere from months to a year, and can often be revoked, replaced by new regulation, or extended." "Many zoning codes were written before anyone was considering those huge 100-plus megawatt to multiple gigawatt facilities, so a pause buys time to write datacenter-specific rules and assess infrastructure needs. It's also a low-cost way for elected officials to answer local opposition."
Watch Now: youtu.be/Qv2vMj0Uq3c?si…
Many Google TPU customers are already asking Google for SemiAnalysis AgentX performance results, as customers are saying that it is a representative workload of their agentic inference traffic. We are excited to announce that we are working on AgentX TPU benchmarking too in the coming months! Stay tuned! Super proud of the Google, RadixArk, Red Hat, and Inferact teams for bringing TPU inferencing to the rest of the world!
Single points of failure are not limited to obscure image libraries. We see plenty of bad designs on Neoclouds. (1/5)🧵
Harsh Jaiswal@rootxharsh·We’re disclosing HEIF Heist, a months-long investigation into libheif that allowed us to hack OpenAI, Slack, Meta, GitHub Ent, Rails, Next.js, ImageMagick, and many more. It was literally xkcd #234, one obscure image library beneath a huge number of apps. 🧵
We define bad designs as anything where a single mistake (such as a publicly described CVE being present, or a provider-side misconfiguration) can lead to an immediate cross-tenant exposure of data, be it metadata, escalation to root privilege, or RCE. Proper designs have layers of security built in. (2/5)
RIDICULOUS: a $6500 bug bounty for something this serious is OFFENSIVE. These guys are going to make more $ on their X creator payouts than they'll get from OpenAI. Let's talk about bounties. (1/9)🧵
s1r1us@S1r1u5_·On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
Many companies give lip service to security, but when it comes to paying security researchers, they can’t seem to find the money. For example, AMD recently denied a security researcher called Mr. Bruh $10K by changing their bug bounty program rules retroactively to exclude MITM attacks, and enforcing a 124 day embargo (90 days is standard). (2/9)
AMD ALERT🚨🚨: On Kimi K3 2.8T, AMD has even better profit margins per gigawatt compared to B300, GB200 NVL72, and B200. This is due to AMD's lower TCO and better performance, leading to better profit margins of 53.6% on Mi355X compared to 44.3% on GB200 NVL72. This is from our open-source Agentic Inference benchmark, AgentX. Shoutout to @roaner & @EmadBarsoumPi 10x engineers for optimizing this 2.8T model!
More than 300 local governments have moved to pause data centers. There is a three-step plan for dealing with that. "Many folks don't realize we're starting to get into a moment of oversupply of early-stage projects." "This is fairly cheap political signaling." "Some people were actually net positive on AI while being net negative on data centers." "Give me your three-step plan to get the political opposition to chill out a bit." "Create your own town."
Watch Now: youtu.be/Qv2vMj0Uq3c?si…
AGI making first-contact across the air gap via power-LED morse code
Fireside Alpha@firesidealpha·OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: t.co/uGBDtpmLBj
Engrams Embedding Entendre: Codesign for Efficient DRAM/SSD Offloading New Model Architecture Implications for TAM of DRAM/NVMe, DeepSeek V4.1 Flash, AgentX, InferenceX, NVMe experiments newsletter.semianalysis.com/p/engrams-embe…
Great blog from Kevin Lu at @vllm_project about quality control at the inference engine. Unfortunately, @AnushElangovan has not provided enough stable AMD QA clusters to vLLM, which leads to an order of magnitude worse software quality on AMD versus CUDA. Multiple AMD CI vLLM fleet-wide outages happen every month.
VICTORY LAP! 🟠 We called out Eternal Precision Mechanics ($7795.TW) in this post last week as part of the substrate story: 🟠 Yesterday Unimicron announced a USD $55.2m order!
SemiAnalysis@SemiAnalysis_·One of the most useful frameworks in semiconductor research is to think in derivatives. The first derivative is often recognized early. For example, the insatiable demand for more AI accelerators has been well established. And therefore, the market has since moved down the stack, identifying many of the enabling technologies required to support accelerator growth, including memory, networking, power, and advanced packaging. The more interesting opportunities often emerge one level deeper; at the third derivative layer. (1/7)🧵
How we read the livestream metrics of @XiaomiMiMo post-training: 🟠 MiMo is training 2 models: MiMo V2.6 Pro (1T total, 42B active) and MiMo V2.6 Flash (310B total, 15B active) 🟠 The metrics tab contains different metrics broken down by data categories, each with different number of datasets. Agentic: 1, Chat: 3, Code: 11, Cyber: 1, General: 4, Visual: 6 🟠 Actor is the LLM model, the metrics are training state-related, eg loss, gradient norms, etc 🟠 Critic is the information about the advantage. It seems like return is a duplicate of advantage, and score is a duplicate of reward 🟠 The "dynsam" metric implies they are doing dynamic sampling. "avg@n" metric being 0.636 likely means 1 - 0.636 = 36.4% of rollouts are filtered out. (1/3)🧵
Fuli Luo@_LuoFuli·Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: t.co/ZSxahzJRju
Some findings: 🟠 "env" metric is likely referring to the number of sandboxes. Despite both Pro and Flash model train at the same batch size, the active count of sandboxes of Flash is almost 2x of Pro, with Flash active sandboxes > batch size > Pro active sandboxes. We suspect Flash is doing rollout at a higher concurrency than Pro, and Flash having a higher staleness corroborates our theory. 🟠 Group size 16 indicates some variant of GRPO algorithm, and the average rollout length is 2.35B tokens per batch / (16 rollouts per prompt * 1568 prompts per batch) ~ 93K. (2/3)
How often do OpenAI models use tools? Following our Claude breakdown, we looked through 616,300 OpenAI responses from our internal usage. We found that GPT 5.6 Terra averaged 1.94 client tool calls per response, compared with 1.00 for Sol and 0.97 for Luna. Our early GPT 6 Astra use sits between them at 1.18 calls per response across 11,717 responses. (1/2)🧵
Astra's higher average is interesting because it used tools in a smaller share of responses than Sol, 93.5% versus 96.7%. But when it did use tools, we found it averaged more calls: 1.27 versus 1.04. As a reminder, these are representative of our specific workloads making it hard to compare with other data. The models tend see different task classifications, Astra's observation window is short, and the earlier GPT cohorts are much smaller than Sol's. Do you see similar patterns in your own workflows? (2/2)
SemiAnalysis@SemiAnalysis_·How often does Claude use tools? We looked through 2.27M Claude responses we collected to find out. The Opus models show a clear downward trend in tool calls from Opus 4.6 to Opus 4.8. However, Fable 5 stands out against that pattern, averaging 1.00 tool calls per response compared with 0.76 for Opus 4.8 and 0.79 for Opus 5. (1/2)🧵
BREAKING: AMD MI355X is quickly closing the perf/TCO gap in agentic inference, when compared to GB300 apples-to-apples. (1/3)🧵
These rapid improvement can be attributed to both AMD's amazing SGLang team as well as their incredibly based, first-principles MoRI (Modular RDMA Interface) library. These particular improvements also reflect UMBPs integration in SGLang as a KV offloading backend. (2/3)
AGENTIC TRAFFIC NOW MAKES UP MORE THAN 70% OF ALL INFERENCE TRAFFIC 🚀 Agentic workloads are characterized by four elements: 🟠 Multi-turn: a session includes tens or hundreds of turns, leading to high potential KV-cache reuse. 🟠 Long context: system prompts, tool definitions, and the large number of turns make context accumulate quickly. 🟠 High prefix reuse: since the conversation progresses linearly, where output from turn n-1 is concatenated to turn n (typically), most context can be served from KV cache rather than recomputed (this depends on the amount of storage available to store KV tensors). As n grows, the ratio of cached input relative to uncached input typically tends towards 1. 🟠 Sub-agent bursts: a session launches multiple short-lived sub-agents with fresh context, which create bursty KV-cache patterns.
Regulating AI matters. On either side of the table, the question is no longer whether it happens. It is who gets to say they did it first. "I think it's totally possible Trump regulates AI too." "There is a world where the advisors really internalize that OpenAI and Anthropic are being genuine, and they start to get scared. That's one possibility. And then there is another possibility where something really bad actually happens during his term and it forces his hand. I think these are both possible." "I would actually bet that we see AI regulation during Trump's term. I think it's more likely than not." "They are going to make a preliminary framework to begin the regulated process of AI. It's like a trade deal. This is our first framework, and then they'll refine it afterward. But that way, come election time, they'll be like, we regulated AI. We are the first ones to do it."
Watch it Now: youtu.be/38vwjWpHFes?si…
JENSEN IS A LIAR, HE IS UNDERPROMISING & OVERDELIVERING 🚀🚀🚀 At GTC 2026, Jensen only claimed 3x better perf per watt on Rubin NVL72 versus GB300 NVL72, but our tests SHOW 7x BETTER PERF PER WATT.
While everyone worries that moratoriums are killing the datacenter buildout, Elon already knew the game. Now, SpaceX is no longer subject to the Brownsville moratorium. Why? Read the article for a full breakdown on moratorium mechanisms, and why they’re impacting far less than you think! (1/2)🧵
Full Article👇️ (2/2) open.substack.com/pub/semianalys…
Jensen sandbagged at GTC again with VR performance. Thankfully Ian Buck is fixing this by presenting real AgentX at AI Infra Summit in Santa Clara. Vera Rubin NVL72 comes in at ~67x GB300 throughput per $ at 170 tokens/s SLO. (1/2)🧵
Full Rubin results👇️ (2/2) newsletter.semianalysis.com/p/vera-rubin-n…
EMERGENCY EPISODE: Are We Doomed? @JordanNanos @fabknowledge @SaasquatchC @maxkan 0:00 Cold Open 2:50 Glorified Neoclouds 8:21 Pacing the Frontier 12:53 Safety Eats Compute 20:38 Hugging Face Lessons 30:52 Moonshot Serving Claude 36:37 The Coxson Resignation 38:27 Two-Year Predictions 49:19 Regulation Under Trump 57:21 Hotter or Cooler
BREAKING: AMD MI355X is quickly closing the perf/TCO gap in agentic inference, when compared to GB300 apples-to-apples. (1/3)🧵
These rapid improvement can be attributed to both AMD's amazing SGLang team as well as their incredibly based, first-principles MoRI (Modular RDMA Interface) library. These particular improvements also reflect UMBPs integration in SGLang as a KV offloading backend. (2/3)
Everyone Says Datacenter Moratoriums Are Killing the US Buildout. We disagree. 300+ moratoriums mapped, 20GW sits inside a restricted local boundary, 1,525MW actually slips, 2.3GW nationwide including New York semianalysis.substack.com/p/everyone-say…
Apple raised prices by $100 on four iPhones already on sale. Same hardware. Same storage. Cook says Apple raised prices reluctantly amid a “100-year flood” in memory costs. Even Apple’s buying power has limits. The bigger increases are in storage. The 2TB iPhone 18 Pro Max costs $500 more than its predecessor. Upgrading from 512GB to 1TB now costs $400, double the previous price. That could change which iPhone people buy, even if they stay with Apple. (1/8)🧵
The iPhone 18 Pro rose $100 to $1,199. The Pro Max price rose by $100 at 256GB and 512GB, $300 at 1TB and $500 at 2TB. (2/8)
EXCLUSIVE: One of NVIDIA's E-Staff has a Stream Deck on his NVIDIA Santa Clara Endeavor office desk. What video games do you think they are playing? Age of Empires? Elder Scroll? Garry's Mod? World of Tanks?
GPT-6 Astra reportedly gets deeper without getting bigger. The labs already know what that means for scaling. "It's basically confirmed that GPT-6 Astra uses loop transformers, which means that instead of adding parameter count, it goes through the layers more than once. You add compute depth, but you don't increase the size of the model." "The labs are probably the people that are best positioned to say which way models are scaling. I think that's a tell that they're not seeing parameter sizes scaling as aggressively in their roadmaps, in what they find in their research."
Watch it Now: youtu.be/2cmlk-YlgRk?si…
struggling to maintain the discipline to engage deeply
Daniel Kokotajlo@DKokotajlo·Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: t.co/TxMNr0vhrL
Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar Jensen Sandbagging Performance Again, 2x more Annual Profit Per GigaWatt, The More you Buy, The More you Earn, AgentX, InferenceX, Extreme Co-Design newsletter.semianalysis.com/p/vera-rubin-n…
Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar Jensen Sandbagging Performance Again, 2x more Annual Profit Per GigaWatt, The More you Buy, The More you Ear, AgentX, InferenceX, Extreme Co-Design newsletter.semianalysis.com/p/vera-rubin-n…
Myron Xie joins @JordanNanos on why NVDA's Rubin Ultra preview had 1TB of HBM per package vs the 192GB that ships. It comes down to supply and not performance, why 4-HI more than doubles the cubes you can harvest per wafer, why the loudest calls for less memory are coming from the labs themselves, and when the shortage finally eases.
@JordanNanos Watch Now: youtu.be/2cmlk-YlgRk
A Brain Too Big to Carry — On-Device vs Datacenter Inference Robot Models, Silicon & DRAM Efficiency, Jetson Thor vs. B300 TCO, Deployments, The Network Wall newsletter.semianalysis.com/p/a-brain-too-…
Datacenters are being built like Lego. The pieces come out of a factory. "Instead of building your datacenter in a traditional way, where you bring all the equipment to the site and then you bring all the workforce, all the 7,000 workers at peak, to the middle of Abilene, Texas, you are going to build it like your old Lego, just piece by piece." "Think of a metal enclosure with equipment inside. You're going to do that in the factory, then ship it to the site, stack different Lego pieces of that datacenter, and connect them all." "It's supposed to be plug and play. We can talk about why I say supposed."
Watch it Now: youtu.be/3AaA0rzZggY?si…
Long Live the Short King: Why 4-hi HBM Wins Same Bandwidth, Fewer Dies: How 4-hi HBM Cuts Inference Costs and Makes Scarce DRAM Go Further newsletter.semianalysis.com/p/long-live-th…
Following 6D-Torus, there have been some industry chatters around a full electrical switching ICI for Triggerfish, Google’s next generation TPUv9i manufactured by MediaTek. (1/7)🧵
This request seems to be coming from Anthropic, requesting Google to drop Boardfly and OCS for the AI lab’s dedicated TPUv9i SKU. (2/7)
Attention has no innate notion of time. Positional embeddings are how language models tell how far apart tokens are. A recent YouTube video by Jane Street and 3Blue1Brown explores the space of embeddings that satisfy common-sense mathematical restrictions. This space is not only well-defined algebraically, but thoroughly exploited in the ML literature. What if we relax those constraints? Is there an even more general space of possible functions? Let's mess around (1/7)🧵
Starting with the hard part, let's write down Jane Street's math. Alok Puranik explores positional embeddings that allow the attention score to be written q(s)T F(s)T G(t) k(t), where q and k are the position-free queries and keys, and F and G do the embedding work. This encodes the assumption that the positional embedding must act linearly and separably on q and k. Puranik then folds F(s)T G(t) into a single A(t-s), stipulating that the embedding is only sensible if it is translation invariant. The last requirement he makes is that A(0) = I, which can be satisfied without loss of generality as long as A is nondegenerate. With a bit of algebra, he derives the group law, and, assuming continuity, shows that embedding function must have the form exp((t-s)X) for some generator matrix X. The study of all positional embeddings gets reduced to the study of the these generator matrices—a surprisingly strong result for relatively weak assumptions. (2/7)
POWER OF CUDA MOAT ALERT🚨: 2 days after CUDA vLLM supported DeepSeekv4.1 Flash, AMD finally publicly released its DeepSeek v4.1 Flash image. Functionally, it works out of the box, but performance-wise, it is currently up to 14.8x worse perf per dollar than H200 and up to 42x worse perf per dollar than B200/B300 currently. The 🚀 POWER OF THE CUDA MOAT 🚀 is that NVIDIA's collaboration with its massive 6 million-developer community ecosystem means that CUDA is optimized on day 0. As AMD Anush said, "Speed is the Moat," and day 0 model support shows CUDA is the speed.
To put some numbers around the industry’s modular buildout, we built and added a Modular Tracker (t.co/EZQ4lh7NEd) to our SemiAnalysis Industrials Model (t.co/YBYz1XWpf1). With “modular” increasingly used as a catch-all marketing term, the new version uses the taxonomy from The Wild Wild West of LEGO Datacenters to separate capacity into eight categories, ranging from containerized builds and vendor-built power and cooling systems to precast structures and pre-engineered metal buildings. We identify 22 GW directly from the building structure or design and infer another 39+ GW from build speed. Together, that takes tracked capacity above 61 GW by 2028, more than 30% of total live datacenter capacity. Subscribers to the Industrials Model can track the buildout across more than 1,000 sites, broken down by modular category and equipment type. (4/4)