
@polynoamial
Researching reasoning @OpenAI | Co-created Libratus/Pluribus superhuman poker AIs, CICERO Diplomacy AI, and OpenAI o-series 🍓 reasoning models
Autocompaction is extremely effective x.com/thsottiaux/sta…
This was one of the bigger open questions in quantum cryptography x.com/42_gravity/sta…
Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss. We’re sharing what we learned from studying a long-running model, and how those findings are shaping our approach to evaluations, alignment, monitoring, and user control. t.co/yePIzJGsAU
2023: LLMs struggle with 4th grade word problems 2024: LLMs can do high school math 2025: LLMs get a gold medal at the IMO Now, GPT-5.6 solves famous frontier math/stat questions. The IMO is today and 5.6 one-shotting a perfect score isn't even news. Where will we be next year? x.com/EdgarDobriban/…
More test-time compute leads to greater intelligence. But as we push ttc from seconds to weeks, latency becomes a bottleneck. GPT-5.6 Sol Ultra scales parallel ttc. The time taken to generate a proof to a 50-year-old problem drops from perhaps a whole day to a single hour. x.com/__eknight__/st…
GPT-5.6 Sol Ultra produced a proof of a 50 year old math conjecture. Unlike the Erdős Unit Distance Problem, this was done with a model publicly available *today*. I look forward to seeing what scientists and researchers are able to do with this model! x.com/__eknight__/st…
There is also a Lean formalization of the proof: x.com/__eknight__/st…
I'm at ICML this week and I'll be doing Q&A today (Tuesday) from 3-4pm at the @OpenAI booth with my reasoning research colleagues. Come by and ask us a question!
Excellent work from @AISecurityInst investigating the impact of test-time compute budgets for frontier AI model evaluations. They make the case even more convincingly than I could! x.com/aisecurityinst…
GPT-5.6 is incredibly strong and fast for coding. I hope we can make it available to everyone soon. x.com/openai/status/…
When we announced @OpenAI o1 some researchers from other labs told me we made a strategic mistake and should have kept it secret so we could accelerate ourselves and pull farther ahead of the competition. Studies like these make me confident we made the right choice. x.com/openai/status/…
Kevin is one of the best journalists covering AI. I especially appreciate how he takes the time to *use* frontier AI and deeply understand its capabilities and limitations. I’m excited to see what he does next! x.com/kevinroose/sta…
I can think of no better person to help shape frontier AI policy than @deanwball. He has a clear understanding of where AI is headed. I look forward to working with him at @OpenAI! x.com/deanwball/stat…
I'm always thrilled to have more Noams at @OpenAI, but I'm especially thrilled to welcome @NoamShazeer! x.com/NoamShazeer/st…
I'm happy GPT-5.5 tops this eval I'm even happier it's still doing the best when measured vs tokens, cost, or wall-clock time! x.com/dawnsongtweets…
We've known about LLM test-time compute scaling since @OpenAI o1. Yet 2 years later labs still report scalar evals for models; safety orgs are still surprised when a scaffold does better via 100x inference; and RSPs still ignore inference budget when deciding critical thresholds. x.com/polynoamial/st…
After AlphaGo, the skill of human Go players noticeably improved. I suspect we will see a similar pattern in math. x.com/wtgowers/statu…
Source: henrikkarlsson.xyz/p/go