
@zacharylipton
Professor: CMU/@acmi_lab, Cofounder: @AbridgeHQ, Creator: @d2l_ai & https://t.co/QQt98VNLUp, Relapsing 🎷
I agree with the premise but not the conclusion. @jackclarkSF’s right that “stochastic parrots”, already a straw man at the time, has aged horribly. But I’m not convinced that the “parrots-infected” crowd would have moved the needle on ASI preparedness. On the technical side, the most powerful labs with the most resources were never “parrot-pilled”. Our poor progress there seems more due to lack of clear direction than to lack of bodies. On the policy side, even today, there’s little consensus on whether to act or what a coherent framework for comprehensive AI regulation might look like.
Jack Clark@jackclarkSF·"stochastic parrot" was a mimetically-fit cognitive virus that spread from 2021-2025; it temporarily blinded many gifted people to the nature of AI progress, burning up crucial years in which they could have helped think through the response to the situation.
I agree with the premise but not the conclusion. @jackclarkSF’s right that “stochastic parrots”, already misguided at the time, has aged horribly. But I’m not convinced that the “parrots-infected” crowd would have moved the needle on ASI preparedness. On the technical side, the most powerful labs with the most resources were never “parrot-pilled”. Our poor progress there seems more due to lack of clear direction than to lack of bodies. On the policy side, even today, there’s little consensus on whether to act or what a coherent framework for comprehensive AI regulation might look like.
Jack Clark@jackclarkSF·"stochastic parrot" was a mimetically-fit cognitive virus that spread from 2021-2025; it temporarily blinded many gifted people to the nature of AI progress, burning up crucial years in which they could have helped think through the response to the situation.
Couldn't be prouder of my student @dkaushik96, as he joins @OpenAI's national security team to work at the forefront of AI policy. DK was a unique student from day 1, and while I tried my best to stats-pill him, he's always had one foot in the technical matter, and another in governance. He's a serious person, a serious force, wrestles with the big questions, and has a big heart. He's open-minded but also sober. Great get for @OpenAI, and great for the rest of us to have someone like him working on the inside. Congrats DK and best wishes for the new chapter!
Divyansh Kaushik@dkaushik96·Some personal news: earlier this week I joined @OpenAI’s national security team. My work has always sat between technical AI, national security, and policy. Over the past year, it has become increasingly clear to me that the hardest questions are moving closer to the frontier itself. 1/
Couldn't be prouder of my student @dkaushik96, as he joins @OpenAI's national security team to work at the forefront of AI policy. DK was a unique student from day 1, and while I tried my best to stats-pill him, he's always had one foot in the technical matter, and another in governance. He's a serious person, a serious force, wrestles with the big questions, and has a big heart. He's open-minded but also sober. Great get for @OpenAI, and great for the rest of us to have someone like him wrestling with the hard questions on the inside. Congrats DK and best wishes for the new chapter!
Divyansh Kaushik@dkaushik96·Some personal news: earlier this week I joined @OpenAI’s national security team. My work has always sat between technical AI, national security, and policy. Over the past year, it has become increasingly clear to me that the hardest questions are moving closer to the frontier itself. 1/
Prior to 2012 just about nobody in the core software industry (MSR doesn’t count) paid much attention to academic CS publishing. Also, prior to 2012, academics actually read papers and judged you more on the quality than the quantity of those papers. Despite the low cost of producing shit papers, even then, the system kinda worked mainly because there was little incentive to spray slop. The groundswell of interest in the outputs of CS academia broke the system. The reward for “has a NeurIPS”, divorced from the “I read this paper and want to interact with this mind” broke the system. Perhaps all that it takes for the system to rebuild is for it to burn to the ground. Once everyone with money and power thoroughly decides to ignore papers, the incentives to spam will evaporate. From the ashes of the scorched landscape of academic publishing a few fresh shoots will sprout. These will be weirdos with ideas that care about their ideas and don’t care if there’s a pot of gold at the end of the review cycle. Open science and academic inquiry will bloom again, until overgrowth and an untimely drought ignite the next wildfire.
Aaron Roth@Aaroth·The cost of producing low quality papers has gone to almost zero, so to maintain the publication system we need to increase the cost of submitting low quality papers. The easiest way to do this is to impose reputational cost for producing low quality work. How to do this?
Prior to 2012 just about nobody the core software industry (MSR doesn’t count) paid much attention to academic CS publishing. Also, prior to 2012, academics actually read papers and judged you more on the quality than the quantity of those papers. Despite the low cost of producing shit papers, even then, the system kinda worked mainly because there was little incentive to spray slop. The groundswell of interest in the outputs of CS academia broke the system. The reward for “has a NeurIPS”, divorced from the “I read this paper and want to interact with this mind” broke the system. Perhaps all that it takes for the system to rebuild is for it to burn to the ground. Once everyone with money and power thoroughly decides to ignore papers, the incentives to spam will evaporate. From the ashes of the scorched landscape of academic publishing a few fresh shoots will sprout. These will be weirdos with ideas that care about their ideas and don’t care if there’s a pot of gold at the end of the review cycle. Open science and academic inquiry will bloom again, until overgrowth and an untimely drought ignite the next wildfire.
Aaron Roth@Aaroth·The cost of producing low quality papers has gone to almost zero, so to maintain the publication system we need to increase the cost of submitting low quality papers. The easiest way to do this is to impose reputational cost for producing low quality work. How to do this?
Apparently Apple has an internal model that folds recursively, keeping it hush to widen lead over Samsung.
Bold prediction: Long after the machines forget about Dario, Sam, and the rest of us entrepreneurs and scientists, they will remember Miles Davis. Who wants to triple dose peptides and sit out WWIII with me in Wyoming so that we can resolve this on what remains of Polymarket in the 30th century?
The era of author credit lasted ~3000 years. How long will the era of prompter-pilot credit last?
What a joy to spend Sunday celebrating the 90th birthday of @yudapearl, a true inspiration and a giant in our field(s). Judea revolutionized AI research twice, first with its probabilistic turn in the 80s (the field had been mired in side quests to develop alternative logics of uncertainty) and again in the 90s/00s with a systematic treatment of causal identification. Judea is a hero for many reasons, not the least of which is his intellectual longevity. His probabilistic turn came in his 40s (and with it the invention of Bayes nets, d-separation, etc) and his second came towards his 60s, including the complete treatment of the do-calculus. So many inspiring people came to pay their tribute, including @DaphneKoller @erichorvitz, with recorded messages sent in by nobel laureates and personal acquaintances alike, all under the careful curation of @eliasbareinboim. Just like @yudapearl, the event could not be contained, with science, tragedy, joy, and raucous political debate sharing the stage (and occasionally wrestling for it). Throughout, I was awed both by Judea’s perseverance (he was once trapped under a bookshelf at 430a because he was still at the office during an earthquake!) and his pursuit of insights so crystalline that they could ripple across the entire canvas of human scholarship. If the development of AI has profited from a generational struggle between the scruffies and the neats, then @yudapearl sits among our greatest neats, and while the scruffies may have the hot hand today, I hope that the next wave of neats (be they humans or machines) will take inspiration from Judea’s life’s work.
The @TransluceAI team has been working tirelessly to understand the mental health impacts of LLM interactions for years now. Awesome to see such thorough work undertaken in the public interest and shared for public benefit.
Transluce@TransluceAI·Today, Transluce is releasing the most expansive independent evaluation to date of how AI systems respond to users experiencing mental health crises. We evaluated 77 model variants from OpenAI, Anthropic, Google, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI.
Judging from their top-of-mind associations for "cult", the author of this article does not live in SF.
Feeling seen.
Vega Shah@dr_alphalyrae·one of my most frustrating personality traits, a vestige from academia, is that when I hear something terrifyingly inaccurate, especially in fields i have studied - I have an urge to rectify. But in truth, for most corporate scenarios it’s better to let some people say the wrong thing
Just a few people exploiting that opportunity would then result in the markets being overall calibrated. Effectively a market-mediated isotonic regression. Key detail here is that calibration says nothing about the strength of the predictions. So charts are wrong to conflate calibration and accuracy.
@LiamFedus (can’t remember which commentators (not you) i saw calling this “accuracy”)
Little in society better demonstrates causality != correlation than Michelin. Stars are highly correlated w great restaurants but Michelin actively makes dining worse. Restaurants sacrifice personality to pursue first star & add ridiculous diaramas to get their second
Thinking some more on a model of the graph. Michelin stars are causally downstream of both “Actual quality” and “Michelin-ness”. Restaurants seeking to gain (or boost) their stars intervene on Michelin-ness, which raises the odds of a star but, regrettably, has a downstream negative impact on actual quality.
History will remember 2026 as either “year zero of the post-human era” - or “the year of the parenthetical hyphen”
These insatiable platform players, storming up the stack to undercut their application layer partners with inferior products at predatory prices. When will it ever end? [i see you @justins]
Breaking news: Brad Lightcap & Denise Dresser joining Barrett Zoph as Co-CFOs at Thinking Machines.
whoever trains a God-tier de-spaghetti refactor code model will make 1T 🫰
We need a new word that rhymes with “moderate” but sounds fucking badass. 1. Unbowed 2. Sovereign 3. Hardcenter 4. … reclaim heterodox? maverick? We need a word that makes @ezraklein sound like an action star.
Less impedance between mind & world. x.com/alpaysh/status…
The only moat left is neurodivergence.
Amazing moment in our Orwellian corporate state that nobody can tell whether @demishassabis stepped up or stepped down.
The only constants anymore are death, taxes, and @GaryMarcus spawning into every AI Twitter war.
First person to name twenty-six 2026 “AI for doing science” startups wins a free trip to Disney World.
Just me or does @ChatGPT search the most random websites no matter the query? If I ask for the best bagel, it will “Searching ” + {arxiv.org, imdb.com, …}
Every doofus with “stealth startup” on their LinkedIn: youtu.be/vnmplSPgwEQ?is…
The History of AI Claims, Compacted Boomers: “We describe a knowledge-based system for…” Gen X: “We prove PAC-theoretic bounds…” Gen Y: “We present a novel architecture…” Gen Z: “We discover a new scaling axis.”
Through the ages, the trendy claim of each era would become so salient that authors would contort their results in bizarre ways in an attempt to license a claim of the desired form.
Derek's thoughtful, and I'm excited to read the whole piece, but got stopped in the first paragraph: It's not actually true that Kimi3 delivers frontier performance while "costing a fraction of the price". The vibes narrative on token economics seem to be getting ahead of reality. Yes, Kimi3 is an amazing model and forces us to radically re-assess our beliefs about the frontier–OSS time gap. Did a little back of the envelope math using @OpenRouter pricing data. Tokens: 40% cheaper that GPT-5.6-Sol... but Kimi3 requires more tokens per prompt, more reasoning tokens, and more output tokens per task. Per task economics: 16% cheaper. If we further account for differential success rate, need for retries / additional calls, probably marginally more expensive. The premise might be directionally right, but seems we're confusing momentum for present reality. On the other hand, if we did this analysis on Kimi3 / Fable the claim absolutely would hold up. And we likely have Kimi & co to thank for the price pressure to drive down the price of Sol and (and now Opus 5).
And here I thought this site couldn't surprise me anymore... x.com/treejordan/sta…
Evergreen, but feeling even greener. x.com/zacharylipton/…
Simple solution for USG: make unambiguously clear that distillation of all kinds is fair game for US-based companies, regardless of ToS. Watch how fast US OSS catches up …
Influencers + AI slop: shameless, unoriginal, gobbled up by X. “That's why the honest line in the launch post is doing real work.” x.com/aakashgupta/st…
\strikethrough{Repost/Quote} Repost/Riposte attn: @elonmusk
Dropping off the grid for a fabulously overdue honeymoon. Shitposts received in my absence will be triaged appropriately & reviewed on my return
I love @LiamFedus. I also love that “back in the early days” now means January. x.com/liamfedus/stat…
Kumon was the original benchmaxxing.
Excited to share to works by my students at #ICML2026 both focused on when and how to leverage synthetic LLM-generated data in statistical analysis, e.g. to lower the variance of estimates from a survey study. Emily’s poster is today at 5p and Pranav’s is today at 3p. x.com/yewonbyun_/sta…
See Pranav’s work which gives an exact characterization of final sample error of estimates leveraging PPI family estimators and gives insights on when they should/shouldn’t be used. x.com/pranavmani30/s…
Don’t worry USA, @realDonaldTrump just made a call to @FIFA and they’re overturning the outcome. We won! 🇺🇸
AI companies complaining about distillation is the single greatest act of hypocrisy in the history of humanity.
“We invested billions and years of effort into imitating knowledge work without permission; and now some upstart companies dare to imitate our work without permission?”
Welcome to 2026, when scientists rely on the religious to provide skepticism. x.com/miles_brundage…
If @ylecun can really make a 10yr old a first-time listener, this will dwarf his contributions to machine learning. Sign me up for the alpha! x.com/ylecun/status/…
Dangerous bioengineering effort with high x-risk averted. Claude dodges another bullet to safeguard humanity.
Is your agent underperforming, or is it your harness? Wrong. Your agent is your harness is your agent. When every (agent-wrapped harness)-wrapped agent meta-(meta-learns), that’s the AI emerging. #feelthelearn
This just in: @JohnJumperSci’s departure followed revelations that he maintained a thinly disguised second career as an SF realtor. #doubleagent x.com/johnjumpersci/…
Drink a shot every time you hear a CEO mumbling about teams of agents.
MXNet undead, resurrected by coding agents. Besides being a fun walk down memory lane, this is cool case study on the behavior & autonomy of coding agents & their capacity to animate legacy code (by @smolix) alex.smola.org/posts/46-mxnet…
What fraction of big lab modeling energy / budget is focused on coding?
People out here abusing hyphens - attn @EffectvAltruism - to convince us they’re not AIs.