
@intology
Automating the process of discovery.
We had a great time presenting at the AI Scientist Summer Workshop in Boston. Thank you to @WengongJin, @YuanqiD, @YEktefaie, @xiwei393, @BotaoYu24, Yikun Zhang for organizing this amazing event, and thank you to @MSFTResearch for hosting us! x.com/intology/statu…
@WengongJin @YuanqiD @YEktefaie @xiwei393 @BotaoYu24 @MSFTResearch And thank you to @AdaFang_ for inviting us and making this all happen.
We will be presenting at the AI Scientist Summer Workshop in Cambridge, this coming Tuesday, August 4, 2026. Ron Arel (@ronusedh) will be presenting. Thank you @AdaFang_ for coordinating! x.com/intology/statu…
Ron Arel (@ronusedh), co-founder of Intology will be on @TBPN today to discuss Locus' automated post-training capabilities, and more x.com/tbpn/status/20…
The models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human post-trained Qwen3 model. Today, LLMs post-trained end-to-end by Locus are in production to millions. 🧵👇 PostTrainBench evaluates agents' ability to post-train models on various domains given 10 H100 hours. We extend PostTrainBench via PostTrainBench+, which has a greatly expanded compute budget that provides clearer signal on automated post-training capabilities. We find that thousands of H100 hours help distinguish methods' performance post-training Qwen3 1.7B-Base models, and that Locus scales best. In this setting, modes trained by Locus collectively surpass the perforamce of the offical human post-trained Qwen3 1.7B model. In a test of generalization, we ran Locus on all live Kaggle competitions with prize money and public leaderboards. After 16 days, Locus achieved the 4th highest average rank among all participants.
Locus is state-of-the-art on PostTrainBench. PostTrainBench gives agents 10 H100-hours to post-train open-weight models on seven benchmarks spanning domains from healthcare to coding. The benchmark uses the performance of the existing human-post-trained models on each of the individual benchmarks as a reference.
Our Artificial Scientist, Locus, achieved a World Record on the NanoGPT speedrun via a fused triton kernel! x.com/intology/statu…
When coding agents consider algorithmic work, they rarely succeed. Instead, they either reason themselves away or regress performance. For example, Autoresearch repeatedly considered reducing the number of value embeddings from 3 to 2, but avoided the change after deeming it risky without any experimentation.
With a high human bar established over a long period of time, built-in contamination prevention, and optimized initialization to reduce the effect of low-hanging fruit, NanoGPT-Bench elicits a clear signal for AI R&D.