
@rsalakhu
CSO @ Sooth Labs, Professor @ CMU, President Elect ICML Board, Ex-VP of Research @ Meta (Multimodal LLMs, AI Agents), ex-Director of AI at @Apple
Looks like my X account has been hacked. Someone has been sending DMs from my account to random people, inviting them to schedule a meeting with me via a Calendly link. Those messages are not from me. Please ignore them. Help me retweet this message. Hey @X, I changed my password but I could really use your help resolving this.
I couldn't agree more. x.com/sriramk/status…
Nvidia 💪! x.com/jensenhuang/st…
Hello I'de like to start my own crypto currency how do i do this
We're officially live on @EasyA_Kickstart This is the beginning to kickstart the production of Sooth Labs, we've partnered with EasyA to launch our first official fundraising round kickstart.easya.io/token/6qTs1aj6…
We are live now! If you'd like to check it out and support the official contract address is: app.uniswap.org/explore/tokens…
Today we’re launching Sooth on EasyA Kickstart (Robinhood Chain). Sooth AI forecasting engines for long-term event prediction and persistent world modeling. Inspired by frontier research in probabilistic models. What major events or trends should we forecast first?
Something important is coming in long-horizon AI forecasting. Frontier research on persistent world models and prediction engines is heading to a new accessible platform via EasyA. Stay tuned.
1) New project: Stratum a proof of work decentralized network for community fine tuning open Kimi weights. The problem is simple: all frontier models are trained behind closed doors. The weights never leave the lab.
2) Stratum is built on a different premise: open weights, forged together. Contributors install a CLI, connect their GPU, and pull LoRA/SFT shard jobs from a coordinator. Each machine trains a small adapter. The coordinator merges shards into public checkpoints you can download from the dashboard. Every user also has a @Solana wallet tied to their identity and get paid rewards in the the $STRAT token that I have made.
I kind of like the new narrative: Open-weight-model-dominant world = full AI communism. It feels a lot safer than the world where Open-weight models = nuclear weapons with humanity's annihilation I think we're making progress. We just need a couple more gradient updates with a big learning rate and we will be fine.
Wait, what? x.com/johnennis/stat…
No, Nitish @nitishsr, the privilege is all mine! Working with you has been an incredible experience, especially as you were one of my very first PhD students. A fun story: During one of our meetings at Apple, an executive asked Nitish: so, what are you an expert in? Another senior IC jumped in: Nitish is an expert at everything. And I can attest to that: Nitish is an expert at everything. Nitish is the co-creator of Dropout and he has amassed nearly 90,000 citations. That is an Impact 💪!
Best times ever! x.com/tang_1c/status…
It feels like just yesterday that you, Yonatan @ybisk, and I were having these long technical discussions about embodied AI, some of my favorite conversations. You'd always come up with truly unique and contrarian ideas, I miss those meetings. @MicrosoftAI is very lucky to have you.
Thanks Hubert @cheehoohubert for the kind words. We should catch up next time I am in California. Fun fact: Hubert’s most cited paper (almost 6,000 citations) is called HuBERT (Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units). His second most cited paper has over 3,000 citations (Multimodal transformer for unaligned multimodal language sequences)! Impact: 💪
Thank you Lisa @rl_agent, for the kind words. You've always had brilliant ideas in RL, and it's been so exciting to see you making an impact on Gemini at GDM. They're very lucky to have you. x.com/rl_agent/statu…
Lol I think Peyman @docmilanfar really likes me. I don't actually know him, but I'm deeply honored by his encyclopedic knowledge of my academic and industry career 😆. But jokes aside, my current and former students at CMU (and Toronto) are absolutely exceptional, self-directed, very independent in shaping and driving their own research agenda, and with impeccable work ethic, much like so many students at CMU. A student's success is also determined not only by their advisor, but also by the environment in which they grow. There is something genuinely special about CMU. It is a very collaborative place with an extraordinary concentration of talented, hardworking, and motivated students. Doing PhD is never easy, it is a lot of very very hard work, especially at a place like CMU. But many of these students will go on to become leaders in academia and industry, professors at top universities, founders, CEOs, and pioneers of new fields. And I will certainly continue to celebrate my students’ achievements, their scientific breakthroughs, and the impact they have on the world.
If you are trying to compare yourself to Kimi, open source the model just like Kimi. x.com/elonmusk/statu…
Charlie @tang_1c was one of my very first PhD students at Toronto. I still remember those days when together with Nitish @nitishsr, we were pitching our research to Apple execs. Great memories and fun times! x.com/tang_1c/status…
I think there's a lot of confusion and misinformation about Zhilin's visa, H1B lottery, immigration, and why he left the U.S. Zhilin had many opportunities to stay in the U.S. if he wanted to. In fact, I was on an email thread with a senior Apple exec (Tim Cook's reports) asking whether Zhilin would consider joining Apple. I said that he wanted to go back to his homeland. The response was, well, we have an office in Beijing if he'd like to join there. But Zhilin was quite determined to go back and build a startup. I remember him telling me that if he didn't at least try starting his own company, he would regret it for the rest of his life. I respect that, and he was right. Many of my international PhD students do choose to stay in the U.S. Of course the U.S. immigration process can be quite intimidating and uncertain, even for superstar PhD graduates from places like CMU.
Paul Liang @pliang279 is the GOAT. If you're thinking about doing a PhD at MIT, I can't imagine a better advisor. x.com/pliang279/stat…
Thank you @mihdalal. Mentoring and learning alongside brilliant students like you has been a true blessing. x.com/mihdalal/statu…
Thank you @dchaplot for kind words. Working with brilliant students such as yourself has been such a privilege. x.com/dchaplot/statu…
Thank you @yisongyue for kind words. x.com/yisongyue/stat…
I’ve been asked several times whether Zhilin Yang, the founder of @Kimi_Moonshot was my PhD student. The answer is yes and he is absolutely brilliant. But I’ve been incredibly fortunate to work with so many outstanding PhD students over the years. So I thought I’d brag a little about them and their career paths (of course there are also many MSc and undergraduate students, sorry if I missed anyone): Founders / Founding Team Members Devendra Chaplot @dchaplot PhD, Founding Member Thinking Machines / Mistral Zhilin Yang PhD, Founder & CEO, Moonshot AI Jimmy Ba @jimmybajimmyba MSc/PhD, Co-founder xAI Hubert Tsai PhD, Co-founder Spuree, Apple Nitish Srivastava @nitishsr PhD, Co-founder Perceptual Machines; Co-founder Vayu Robotics Charlie Tang PhD, Co-founder Perceptual Machines, DE Shaw. Professors Paul Liang @pliang279 PhD, MIT Ben Eysenbach @ben_eysenbach PhD, Princeton University Ruosong Wang @RuosongW PhD, Peking University Bhuwan Dhingra @bhuwandhingra PhD, Duke University Roger Grosse @RogerGrosse Postdoc, University of Toronto Alexander Schwing Postdoc, UIUC Research Scientists Shuyan Zhou @shuyanzh36, Postdoc, Meta Superintelligence Lab Tiffany Min @SoYeonTiffMin PhD, Microsoft AI Murtaza Dalal @mihdalal PhD, Tesla AI Minji Yoon @MinjiYoon90 , PhD, Microsoft AI Shrimai Prabhumoye PhD, NVIDIA AI, Mistral Haitian Sun @sun_haitian PhD, Google DeepMind Emilio Parisotto PhD, Google DeepMind Lisa Lee PhD @rl_agent, Google DeepMind Manzil Zaheer @ManzilZaheer PhD, Google DeepMind Jamie Kiros PhD, Google Brain, OpenAI Yuri Burda PhD, OpenAI, Anthropic Cody Severinski PhD, Amazon
I agree. x.com/pmddomingos/st…
Thank you for kind words. x.com/AbdoDameen/sta…
Congratulations to Zhilin Yang, founder and CEO of @Kimi_Moonshot, on the latest Kimi release. What a huge win for the open-source community! It feels like just yesterday Zhilin was graduating from my lab at CMU, jointly co-advised with William Cohen. Not only did he complete his Ph.D. in just four years, but he also made truly fundamental contributions to ML during his time at CMU. What a spectacular career! Congrats again Zhilin, and thank you and the entire Kimi team for everything you're doing for the open-source community.
Super cool effort. Archer Announces Zee, AI Foundation Model Purpose-Built for Aviation, a Key Pillar of Its Physical AI Strategy news.archer.com/archer-announc…
Pre-Training Isn’t Bitter Enough – Machine Learning Blog | ML@CMU | Carnegie Mellon University blog.ml.cmu.edu/2026/06/17/pre…
New paper on Understanding Compositional Generalization in Language Model Reasoning Paper: t.co/pUIPP7vZ02 This work studies why combining SFT with RL is so effective for post-training reasoning models. We argue that its success stems from compositional generalization, formalized through a hierarchical latent selection model in which reasoning is built by composing reusable atomic modules. We show theoretically that SFT and RL play complementary but asymmetric roles: SFT provides compositional reasoning traces containing reusable atomic modules, while RL discovers and recombines these modules to achieve compositional generalization. Experiments show that the strongest performance is achieved when SFT provides broad coverage and RL explores compositions beyond the SFT distribution. See a more detailed thread by @LingjingKong.
Congratulations to @GoogleDeepMind Computer Use Agents (CUA) team for taking the #1 spot on Odysseys with a vision-only (GUI-based) agent using by Gemini 3.5 Flash. Impressive result given the difficulty of Odysseys and the fact that it is a Flash model. t.co/rj5BHK5g6C Odysseys evaluates realistic, multi-hour web workflows that require sustained planning, memory, reasoning, and verification across numerous websites and tools, beyond short single-step browser tasks. Exciting to see progress toward truly capable long-horizon web agents.
Exciting work at MPG Ranch leveraging multimodal data collection, integrating ground-level observations with aerial imagery to build a richer understanding of ecological systems.
iOSWorld: A Benchmark for Personally Intelligent Phone Agents Paper: arxiv.org/abs/2606.09764 Code+Environment: iosworld.io x.com/rsalakhu/statu…
True story. An amazing project. And one that I think will be quite impactful. x.com/JangLawrenceK/…
Congrats to the @browser_use team for taking the #1 spot on Odysseys, a highly challenging benchmark for long-horizon web agents: t.co/rj5BHK5NWa Odysseys evaluates realistic, multi-hour web workflows that require sustained planning, memory, reasoning, and verification across numerous websites and tools, far beyond short single-step browser tasks. Exciting progress toward truly capable long-horizon agents.
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents Paper: t.co/uoA0hEXqUe Code, Environment and Tasks, Agentic harness: t.co/4VT2oZ9Pcp MyPCBench provides a Linux desktop environment with 17 simulated real-world web applications and a complete desktop software stack. The benchmark contains 184 tasks inspired by authentic OpenClaw community requests and evaluates agents through a unified computer-use + bash interface. We benchmark leading closed- and open-weight models and find that the strongest model, Claude Opus 4.6, successfully completes 55.4%of tasks. We find that failures are concentrated on long-horizon, multi-application workflows, highlighting that personalization and persistent user context remain key challenges for the next generation of AI personal assistants. Joint work with @JangLawrenceK, @andrewkjang7 and @kohjingyu.
New work on FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning: Paper: t.co/7PXIUrXpAK Web: t.co/IeJeTJ27L8 FACTR 2 shows that learned force signals can both enable force-feedback teleoperation on low-cost manipulators and improve behavior cloning (BC) policies for contact-rich tasks. It consists of two components: 1. Neural External Torque Estimation (NEXT): A lightweight model that infers external joint torques without dedicated force sensors. 2. Force-Informed Re-Sampling Training (FIRST): A training strategy that uses the learned force signal to identify and upsample task-critical moments. The key insight is simple: policy failures rarely occur in free space, they occur during brief pre-contact alignment and contact-rich interactions, where precise corrections matter most. Together, NEXT and FIRST bring force-aware teleoperation and robust long-horizon contact-rich policy learning to off-the-shelf robot arms, without requiring additional sensing hardware. See a more detailed thread by @JasonJZLiu.
New work: iOSWorld: A Benchmark for Personally Intelligent Phone Agents Paper: t.co/d2mtBEjcMu Code+Web: t.co/gxlt6wFwPL An interactive benchmark built around a persistent user identity spanning 26 custom iOS apps, including analogs of OpenTable, Uber, DoorDash, AirBnB, Chase. These apps contain richly interconnected data, such as messages, transactions, travel histories, social relationships, and personal preferences. iOSWorld comprises 133 tasks across three levels of difficulty: single-app (27), multi-app (60), and memory & personalization (46). We evaluate leading frontier and open-source computer-use agents under both vision-only and privileged vision+XML settings. Even with privileged access, the strongest frontier model achieves only 52% success, underscoring the challenge of personalized, cross-app reasoning. We also release an MCP-based tool-use interface for all 26 apps, enabling controlled comparisons between computer-use, tool-use, and hybrid agents. The full benchmark includes the apps, seeded user data, tasks, rubrics, evaluation code, MCP server, and cloud-based infrastructure for running experiments without Mac hardware. See a more detailed thread by @JangLawrenceK.
New work on Multi-Agent Computer Use (MACU). The future of computer-use agents lies in multi-agent systems that combine planning, coordination, and parallel execution. Paper: t.co/6rkHjTR9J0 Webside + Code: t.co/Ng7Bwz3feh MACU introduces a manager agent that decomposes tasks into a dynamic directed acyclic graph (DAG) of subtasks, dispatches parallel subagents, and continuously updates the plan as new information arrives. Across OSWorld, Online-Mind2Web, WebTrailBench, and Odysseys, we see performance improvement by 4.7–25.5%, achieving better test-time scaling, and solving long-horizon tasks that single-agent systems often fail to complete. On Odysseys, MACU reduces task completion time by 1.5×, showing that multi-agent coordination is a powerful path toward more capable and efficient computer-use agents. See a more detail thread by @kohjingyu.