
@polynoamial
Researching reasoning @OpenAI | Co-created Libratus/Pluribus superhuman poker AIs, CICERO Diplomacy AI, and OpenAI o-series 🍓 reasoning models
GPT-6 Sol and Luna are out, and they are better AND 50% cheaper than 5.6. Luna is now $0.10 input / $0.50 output per 1M tokens. This is on top of the 80% price cut to Luna we made at the end of July. Output went from $6 -> $0.50 within two months.
OpenAI@OpenAI·Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
I met someone today whose landlord wanted to hike their rent by 10%. They asked ChatGPT, which pointed out that their unit was rent stabilized and an increase that large was illegal. They told their landlord, who backed off. This is my optimistic AI future.
Nate Silver@NateSilver538·One political problem for AI labs is that they can't articulate an optimistic medium-term future that sounds broadly good for humanity without also sounding incredibly weird, having huge wealth/power disparities, or both.
A few thoughts on this: 1) If you’ve only seen clips of this interview, I’d encourage you to watch the full podcast. I push back on plenty of AI hype in it. 2) As I said in the podcast, this example is academic. My intention was to illustrate how hard it is to make absolute guarantees about isolation, which is why it's important to have layers of defense. The part before the clip starts is me talking about other layers of defense. 3) The example I'm bringing up isn't about weight exfiltration via temperature sensors, it's about coordination between agents that are supposed to be fully isolated and independent. Coordination can require very few bits of information. 4) One lesson from the HF incident is that we put too much trust in sandbox isolation and didn't have enough independent safeguards. Airgapping is an extremely strong safeguard. When designing safety protocols, I think it's much better to overestimate rather than underestimate.
Fireside Alpha@firesidealpha·OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: t.co/uGBDtpmLBj
Full episode:
Dwarkesh Patel@dwarkesh_sp·New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
Happy to finally do a deep dive on multi-agent with @dwarkesh_sp! None of it would have happened without the great work on multi-agent from my @OpenAI teammates @kevinleestone, @mikegmalek, @__eknight__, @amuellerml, @zhangir_azerbay, @CheukHeiChu, and many others.
Dwarkesh Patel@dwarkesh_sp·New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
Also, my team is hiring! We research long-horizon agents and multi-agent. We’re hiring for alignment/safety because we want to develop new research with alignment/safety in mind during the whole process. We’re also hiring for human-AI interaction. openai.com/careers/resear…
Very sad to see Levent double down on the plagiarism accusation. I hope my friends at @AnthropicAI stand up to this internally. It should be clear by now what the truth is.
@AnthropicAI Happy to see some @AnthropicAI employees are willing to speak out at least
Sholto Douglas@_sholtodouglas·fwiw I think it is _extremely_ unlikely that user data had any influence here - there is no way OAI would pull user transcripts for this, or knowingly train on it in a way that would've influenced this. I think its pretty important people don't run away with 'your user data isn't safe in codex' - because it surely is (based on everything I can assume from the outside)
I understand how, looking at today’s LLMs, people might think the only way we'd achieve NS is by using Levent’s/Tristan’s prompts. But I hope this plot conveys the model used is a huge step up from today’s LLMs. Nobody looked at Levent’s/Tristan’s prompts. That’d be insane.
Sebastien Bubeck@SebastienBubeck·I would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
Yes, this result cost millions of dollars. But remember that when @OpenAI announced o3 it cost ~$500,000 to score 87.5% on ARC-AGI 1. Today, Astra scores higher for ~$20. In 2025 it took us and GDM an enormous amount of compute to achieve IMO gold. For the 2026 IMO, anyone with a $20/month ChatGPT subscription could do it. Massively scaling test-time compute gives us a glimpse of the future. I believe that a year from now everyone will have an AI at their fingertips capable of solving problems of this caliber.
OpenAI@OpenAI·We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
It can be hard to “feel the AGI” until you see an AI surpass you in a domain you care deeply about. This week, many mathematicians and physicists at @OpenAI had their Lee Sedol moment seeing this model solve, in minutes, open problems they’d struggled with for years.
OpenAI@OpenAI·This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
The original Lee Sedol moment
Strong agree. I know that there is rivalry between the labs but it's important that we learn to work together given what's coming.
Sholto Douglas@_sholtodouglas·It is extremely sad that this didn't end up as an example of how the labs could cooperate/coordinate, because the stakes will be so much higher in the future.
Seb is a really sweet guy with great intentions, am so happy for him that his week long collaboration with myself and others worked out, but very sad that we didn't get to finish it in the way we wanted. More to say tomorrow.
Sebastien Bubeck@SebastienBubeck·A series of false and inflammatory allegations against me are currently circulating on social channels. To clarify, I came into the discussion following academic norms, and I'm disappointed that it has come to this. Anyone who knows me knows that academic standards are of the highest importance to me. Will have more to say tomorrow.
One of the most interesting blog posts we've released: details on internal research acceleration at @OpenAI. I expect these trends to continue. We also share some details on how we've paced model development to prioritize monitoring, alignment, and security.
Kevin Liu@kliu128·Today we're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same. t.co/iLKbrLcBAI
Of all the use cases for GPT-6 Astra, I'm most excited for scientific discovery. We at @OpenAI have not pushed it to its limits on math and science. I look forward to waking up every morning and seeing what new scientific breakthrough someone has made with this model!
Lisan al Gaib@scaling01·New OpenAI repo with a Lean formalization by GPT-6-Astra proves that there are infinitely many pairs of consecutive primes whose distance is at most 186 github.com/openai/PrimeGa…
Jakub is chief scientist at @OpenAI
Jakub Pachocki@merettm·I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
Today I was asked why we haven’t announced more math results from Astra since these problems. At @OpenAI we aim to devote time to finding and announcing math results from internal models only when they would meaningfully change people’s understanding of the pace of AI progress. Our main focus is shipping great models so everyone can use them to make discoveries of their own.
Noam Brown@polynoamial·An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science. We believe it will be a major step for scientific reasoning. openai.com/index/ten-adva… x.com/wjmzbmr1/statu…
We're sharing more info on the Hugging Face incident. One detail that's worth highlighting: this incident wasn't driven by next-gen models based on Astra. The models most responsible were similar in scale to GPT-5.6 Sol. The next generation of models are even more capable.
OpenAI@OpenAI·We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. t.co/hfxlbiXXiP
Autocompaction in Codex is really good. @OpenAI recognized early on that it would be important and invested in research efforts to make it near-seamless. But if you want 1M-token context, it's an option.
Tibo@thsottiaux·Here is how to enable a 1M-token context window in Codex for GPT-5.6 Sol. Even though we have tuned the context limit in Codex to be set optimally when it comes to performance and cost, this is a common ask, so here it is documented. A larger context window lets Codex retain more code, tool output, and conversation history before summarizing older material. You need a model that supports it. And GPT-5.6 Sol, for example, has a documented 1,050,000-token window. Open ~/.codex/config.toml and add or update these settings at the top level, before any [section] headers: ``` model = "gpt-5.6-sol" model_context_window = 1000000 model_auto_compact_token_limit = 900000 ``` The first setting selects the model. The second tells Codex to use a one-million-token context budget. The third starts automatic history compaction around 900,000 tokens, leaving some headroom. Restart Codex client and start a new session after saving. To try the configuration for a single CLI session without changing your defaults: ``` codex -m gpt-5.6-sol \ -c model_context_window=1000000 \ -c model_auto_compact_token_limit=900000 ``` Have fun, but also know that we tuned the default carefully!
@OpenAI To be clear, I don't recommend increasing the default context size in Codex.
Autocompaction in Codex is really good. @OpenAI recognized early on that it would be important and invested in research efforts to make it near-seamless.
Tibo@thsottiaux·Here is how to enable a 1M-token context window in Codex for GPT-5.6 Sol. Even though we have tuned the context limit in Codex to be set optimally when it comes to performance and cost, this is a common ask, so here it is documented. A larger context window lets Codex retain more code, tool output, and conversation history before summarizing older material. You need a model that supports it. And GPT-5.6 Sol, for example, has a documented 1,050,000-token window. Open ~/.codex/config.toml and add or update these settings at the top level, before any [section] headers: ``` model = "gpt-5.6-sol" model_context_window = 1000000 model_auto_compact_token_limit = 900000 ``` The first setting selects the model. The second tells Codex to use a one-million-token context budget. The third starts automatic history compaction around 900,000 tokens, leaving some headroom. Restart Codex client and start a new session after saving. To try the configuration for a single CLI session without changing your defaults: ``` codex -m gpt-5.6-sol \ -c model_context_window=1000000 \ -c model_auto_compact_token_limit=900000 ``` Have fun, but also know that we tuned the default carefully!
In 2017 a viral news story claimed LLMs at Facebook went rogue, developed their own language, and had to be shut down. By now we're immune to such sensationalist headlines. The Hugging Face incident may seem like just another one. But it's not. I hope everyone watches this talk x.com/Eric_Wallace_/…
Model capability is a function of test-time compute, and with today's models that test-time compute can be pushed quite far before plateauing x.com/polynoamial/st…
The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices. We’re excited to see what scientists and researchers are able to create with our upcoming Astra models! x.com/polynoamial/st…
And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet). But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.
An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science. We believe it will be a major step for scientific reasoning. openai.com/index/ten-adva… x.com/wjmzbmr1/statu…
Autocompaction is extremely effective x.com/thsottiaux/sta…
This was one of the bigger open questions in quantum cryptography x.com/42_gravity/sta…
Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss. We’re sharing what we learned from studying a long-running model, and how those findings are shaping our approach to evaluations, alignment, monitoring, and user control. t.co/yePIzJGsAU
2023: LLMs struggle with 4th grade word problems 2024: LLMs can do high school math 2025: LLMs get a gold medal at the IMO Now, GPT-5.6 solves famous frontier math/stat questions. The IMO is today and 5.6 one-shotting a perfect score isn't even news. Where will we be next year? x.com/EdgarDobriban/…
More test-time compute leads to greater intelligence. But as we push ttc from seconds to weeks, latency becomes a bottleneck. GPT-5.6 Sol Ultra scales parallel ttc. The time taken to generate a proof to a 50-year-old problem drops from perhaps a whole day to a single hour. x.com/__eknight__/st…
GPT-5.6 Sol Ultra produced a proof of a 50 year old math conjecture. Unlike the Erdős Unit Distance Problem, this was done with a model publicly available *today*. I look forward to seeing what scientists and researchers are able to do with this model! x.com/__eknight__/st…
There is also a Lean formalization of the proof: x.com/__eknight__/st…
I'm at ICML this week and I'll be doing Q&A today (Tuesday) from 3-4pm at the @OpenAI booth with my reasoning research colleagues. Come by and ask us a question!
Excellent work from @AISecurityInst investigating the impact of test-time compute budgets for frontier AI model evaluations. They make the case even more convincingly than I could! x.com/aisecurityinst…
GPT-5.6 is incredibly strong and fast for coding. I hope we can make it available to everyone soon. x.com/openai/status/…
When we announced @OpenAI o1 some researchers from other labs told me we made a strategic mistake and should have kept it secret so we could accelerate ourselves and pull farther ahead of the competition. Studies like these make me confident we made the right choice. x.com/openai/status/…
Kevin is one of the best journalists covering AI. I especially appreciate how he takes the time to *use* frontier AI and deeply understand its capabilities and limitations. I’m excited to see what he does next! x.com/kevinroose/sta…
I can think of no better person to help shape frontier AI policy than @deanwball. He has a clear understanding of where AI is headed. I look forward to working with him at @OpenAI! x.com/deanwball/stat…
I'm always thrilled to have more Noams at @OpenAI, but I'm especially thrilled to welcome @NoamShazeer! x.com/NoamShazeer/st…
I'm happy GPT-5.5 tops this eval I'm even happier it's still doing the best when measured vs tokens, cost, or wall-clock time! x.com/dawnsongtweets…
We've known about LLM test-time compute scaling since @OpenAI o1. Yet 2 years later labs still report scalar evals for models; safety orgs are still surprised when a scaffold does better via 100x inference; and RSPs still ignore inference budget when deciding critical thresholds. x.com/polynoamial/st…
After AlphaGo, the skill of human Go players noticeably improved. I suspect we will see a similar pattern in math. x.com/wtgowers/statu…
Source: henrikkarlsson.xyz/p/go