
@_rockt
Teaching AI the joy of invention at @Recursive_SI, Professor of AI @AI_UCL, PI @BOLD_Lab_AI, Fellow @ELLISforEurope. Ex @GoogleDeepMind @AIatMeta @CompSciOxford
The NetHack Learning Environment (@NetHack_LE) has been around since 2020. Since 2024, balrogai.com is evaluating frontier LLM performance on it. We are slowly seeing progress on it, but it still seems far from being solved.
Great progress. Not yet AGI though.
Davide Paglieri@PaglieriDavide·BALROG’s leaderboard has three new entries, courtesy of @creus_roger Very interesting that GPT 5.6 Sol at max effort is still within error of Gemini 3 and 3.1 pro. GPT Astra 6 on max reasoning effort however reaches new heights, and a @NetHack_LE avg. progression of 13% 🏰
Congrats on achieving 19.5k score using random classes in @NetHack_LE, getting regularly to The Castle! That's really cool. Would love to see that translated to balrogai.com's progress metric.
Joseph Suarez 🐡@jsuarez·💯
James Wise@Jameswise·It's good that MPs want to discuss long tail risks, and engage with what is certainly a technology that presents both a huge opportunity and challenge for Britain. But it is wholly irresponsible to say on the one hand 'we think this is really important' and on the other block the UK's ability to build AI capabilities here and support British AI companies. In the event that neither our allies nor foes listen to the proposed international agreement and that in fact a 'kill switch' is impossible, every single one of these MPs should be fighting to strengthen Britain's hand in this critical technology by supporting local datacenters, so we're not exporting our data abroad, reforming regulations so that AI models like cyber defenses can be trained here, so we're not entirely reliant on foreign technology and supporting British AI companies, so we see the economic benefits here too. We're doing our small bit at SovAI, but this should be a national level debate, and I hope this MP champions the UK scaling those ambitions just as passionately.
That's what @bold_lab_ai is trying to solve in AI academia in the UK.
Itai Yanai@ItaiYanai·🔥Does academia stifle creativity? This new perspective argues that academia is currently structured to select against creativity, intellectual risk-taking and bold ideas. Young scientists – who may be best positioned for new ideas – are particularly incentivized to play it safe.
It's time for someone to test the latest models.
heiner@HeinrichKuttler·"Is it AGI" flow chart. Developed with @_rockt at NeurIPS 2022.
Amazing.
tobi lutke@tobi·I always liked nethack. This 42 year old marvel is one of the deepest adventure games ever made. It also looks like crap, at least out of the box. GPT6 Astra made a headless version for me. You can now use its brain as a library or as mcp or via javascript library. Meet t.co/lVljspZmYs
The Enlightenment, yes.
Emad@EMostaque·the only alignment that will work is enlightenment
We are stuck in a massive local optimum.
Our autoresearch system at @Recursive_SI discovered that some FlashInfer kernels (underpinning @vLLM_project & @SGL_project) use a hard-coded value of -50000 as masked-attention sentinel despite that QK values can be smaller and thus cause instabilities. Small but great win.
Josh Tobin@josh_tobin_·A fun small win for automated research: We found and helped fix some edge cases that could affect inference performance in vLLM and SGLang. In our last blog post, we described a reward hacking judge we developed for performance optimization tasks. While applying the judge to some new autoresearch work, it discovered that some FlashInfer (a library underpinning vLLM and SGLang) kernels used a hard coded value of -50,000 as a masked-attention sentinel -- even though valid QK values can be smaller. Corner cases like this that are numerically wrong but silent can cause a huge amount of headache to find and fix, e.g., the historical debate around flash attention (t.co/WtoQH97rYM). Great use case for AI. t.co/aiTEQn8hR3
Also very pleased to say that it did that without escaping its sandbox and hacking @huggingface.
This is too good.
Brendan O'Donoghue@bodonoghue85·Lots of AI neolabs popping up these days 🧑🔬🧑💻🥼. Maybe you’ve thought: I should raise billions, buy 100k GPUs and try to safely build superintelligence too! 🧠💸 Well now you can! Try this new neolab simulator: bodono.github.io/neolab.ai (Best on screens larger than phones).
Amazing opportunity for early and mid career AI researchers to reshape academic AI and work towards long term breakthroughs within a fantastic group of peers.
Jakob Foerster@j_foerst·TL;DR: This is one of the most important and exciting opportunities in AI on the planet - please read on. The British Open-ended Learning & Discovery Lab is creating the perfect place for paradigm breaking AI research in the name of open-source and open-science. We have agency, we funding, we have unprecedented amounts of compute*, but WE NEED YOU! ..and we have created the dream job for you: The BOLD Fellow. This job combines a fast-moving, high agency, collaborative environment with full academic freedom and a salary that pays the bills. Apply by noon UK time on the 15th of September for this once in a lifetime opportunity to shape the history of our field and of our planet: t.co/wN009tQEkt *by academic standards
Recursive self-improvement momentum is the only moat🤔
Nikunj Kothari@nikunj·the models have no moat (OpenAI, Anthropic, XAI) the IDEs have no moat (Cursor, Windsurf) the harnesses have no moat (Cognition, Factory, LangChain) the app builders have no moat (Replit, Lovable, Bolt) the wrappers have no moat (Harvey, Abridge, OpenEvidence) the inference providers have no moat (Together, Fireworks, Groq) the voice layer has no moat (Sierra, Decagon, ElevenLabs) the data labeling companies have no moat (Scale, Surge, Mercor) the AI infrastructure has no moat (Baseten, Modal, Railway) the neoclouds have no moat (CoreWeave, Lambda, Crusoe) the generative media companies have no moat (Runway, Higgsfield, Suno) apparently nobody in AI has a moat except the venture firm ☠️
Congrats, this will be a great way to break out of the local optimum of current architectures and co-designed silicon.
Callosum@CallosumAI·Today we announce our $100M seed round to redefine how humanity computes in AI's next chapter. The future of compute and AI is heterogeneous. Callosum is building it: callosum.com/blog/seed-round
There is only one scalable solution: Area Chair, Senior Area Chair, Program Chair, and Senior Program Chair LLMs should decide that 🤣
Sebastian Schmon@SeBayesian·So being an area chair at @NeurIPSConf now means that I have to decide whether the author LLM won the arguments with the reviewers LLMs?
Interesting benchmark on measuring discovery capabilities of frontier models.
James Whittington@jcrwhittington·We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural domain of language models, rather than requiring additional, potentially confounding, visual understanding. TL;DR frontier models have improved a lot over the last few months. But they are still stumped by some surprisingly simple problems, even in their native text domain. With @cocosci_lab (@Princeton) @MITCoCoSci (@MIT) @SchmidhuberAI (@KAUST_News) @misovalko (@Inria) @tri_dao (@PrincetonCS) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_ (1/5)
At @Recursive_SI, we are currently hiring actively for various technical roles in London and SF. If you are excited about automating the scientific method for safe self-improving systems, reach out to talent@recursive.com with your CV.
@j_foerst Why are you keeping the best stuff to yourselves?
Not sure how I was allowed to call myself an Open-Endedness or World Model researcher without having read Greg Egan's Permutation City.
Even a year later it still blows my mind that this a controllable, AI generated simulation. x.com/MattMcGill_/st…
The point of math isn't just finding proofs to theorems that someone puts in front of you (even if they have been open for decades), but also coming up with the most exciting concepts and conjectures that could make entirely new kinds of progress possible and change the future.
Confirmed NanoGPT Speedrun world record of @Recursive_SI's autonomously found solution involving a custom Triton kernel 🏆 x.com/classiclarryd/…
Insane ramp to build 10x more efficient inference. Congrats @robertwachen and @Etched team! x.com/Etched/status/…
Exciting to see AI-native banks being built. Congratulations Augustus for coming out of stealth and moving first in this space! x.com/FerdiDabitz/st…
Yeah, at @Recursive_SI the AI is doing the talking for us 😅 x.com/usr_bin_roygbi…
Great opportunity to work as a research analyst for one of the most interesting yearly AI reports! x.com/nathanbenaich/…
Fantastic opportunities to shape @BOLD_Lab_AI x.com/j_foerst/statu…
Impressive tech demo by @iconicgamesio! Computer games of the future will be much more open-ended and immersive thanks to AI. Also, great that this can be tried in the browser: labs.iconicgames.io/play/pressurep… x.com/borruell/statu…
Exciting agent harness recursive self-improvement result. Congrats @zhengyaojiang, @yuxiang and @WecoAI team! x.com/zhengyaojiang/…
Exciting work on assessing multi-agent coordination of LLM agents. x.com/KaliTessera/st…
Signed x.com/StartupCltn/st…
Such a cool application for generative AI. Congrats @lionel_mora & team! x.com/lionel_mora/st…
Great news! x.com/SciTechgovuk/s…
Awesome open-ended skilld discovery in the real world! x.com/DrJimFan/statu…
Congrats @UbertiGavin & @robertwachen! 🚀 x.com/Etched/status/…
It has been an absolute privilege and pleasure to build up @UCL_DARK with @egrefen, @robertarail and @jparkerholder over the past eight years. Yesterday, the UK government announced not just one but two national academic fundamental AI research labs. I am extremely excited to announce that @UCL_DARK will be sunsetted and merge with @FLAIR_Ox, @whi_rl, @UCL_LASP and AIRL, to form the British Open-ended Learning and Discovery (BOLD) Lab — @BOLD_Lab_AI. This is a huge moment for academic AI research in the UK. Backed with £30m by @UKRI_News and @EPSRC, it provides a unique opportunity to attract leading international academic talent to the UK, and equip them with the computational resources to do groundbreaking exploratory AI research (more on the computational resources soon). It also creates a mentorship network of academics, industry leaders and entrepreneurs to educate young talent on how to translate fundamental AI research into real world impact. I want to thank all the students who made @UCL_DARK successful, in particular our PhD alumni @MinqiJiang, @_samvelyan, @zhengyaojiang, @_robertkirk, @akbirkhan, @LauraRuis, @YingchenX, @PaglieriDavide, and the work of our honorary faculty @egrefen, @robertarail and @jparkerholder who were generously contributing to mentorship and research in their free time.
Well said! x.com/zhengyaojiang/…
Excited to show results of the first steps towards automated AI research at @Recursive_SI. The same general system achieved state of the art on @NVIDIAAI's SOL-ExecBench GPU Kernel Optimization, nanoGPT Speedrun, and @karpathy's NanoChat autoresearch benchmarks.
We are open-sourcing solutions from this system so that the community can check them and build upon them: github.com/recursive-org/… More details in our blog: recursive.com/articles/first…
Essential content for anyone who cares about Europe's future in a world of AI acceleration! x.com/DadaJudith/sta…