
@MillionInt
CEO and co-founder of @coreauto former VP of RL @ OpenAI : reasoning models, o3, o1, GPT4, ChatGPT, Codex, RL for robots cautious AI optimist
The road to singularity will be paved with blackberries of our generation. It is easier than ever to achieve a glimpse of growth and success but building a long term viable company is as hard as it ever was.
This was once revealed to me in a dream
More questions are deep learning questions than people realized. We just stopped asking deep learning questions and learned to work around them by tuning hyperparameters.
We may have done it at @coreauto but the model wasn't very good at math
Rosmine@rosmine·We need to be dog maxxing our models There is 0% chance of doom if the chain of thought is user: hello model: <think> Oh my god it's a human. I love humans. I will say hi to him, maybe he will give me a problem to chase. I love chasing problems. I will chase the problem for him then he will love me and I will be his best friend. Must help human </think> model: hello! what can I help you with today?
Internet may collapse if GPT-7 is so good it understands all the vagueposts
Inside every lab are two wolves: “I don’t want to destroy the world with my invention.” “I don’t want to lose revenue to other labs I don’t like.” The one you feed is the one that wins.
This is very cool
rohan anil@_arohan_·While we don't have the right solutions for everything, I have a very practical proposal. I don't think open source should be banned - it does sound dystopian and anti freedom in a way that's hard to stomach. I continue to have doubts about economics of open source models, but when there are participants in the market who open source their models we should celebrate that. What we should care about, is that aligned models have much more compute behind them than misaligned models. In the end this will be the blockchain security model that will keep our civilization afloat. I think big companies behind AI development have the right incentives and will try to do their part here. But there will be tons of smaller players fine-tuning open source models on tons of different objectives. An arrangement, where companies actively developing cutting edge AI agree to develop and share highest quality pro-alignment environments with the world can be a huge boon. We don't need to share models, we don't need to share compute. But we need to share values so that there are more good models in the world than harmful ones. Finetuning models on bad, low quality environments is actively harmful and leads to reward hacking. If anyone fine-tuning the models can with low effort align them to shared pro-prosperity and pro-democratic values, we likely have won as a civilization.
Jerry Tworek@MillionInt·Alignment is not as hard to solve as many claim, but it is in the end an algorithmic problem which a lot of ML community mostly stopped working on. Formulation of alignment stated by three (or four) laws of robotics can take us very far, so we roughly know the objective. The tricky part is, how do we take gradient with respect to alignment? We have two algorithms right now at our disposal: pretraining and RL. Pretraining takes gradient with respect to next token prediction - that's def not the alignment objective. We could create RL environments that embody the alignment objective, but: - those are expensive to create, so often cheaper, hackable proxies are used in practice - RL as an objective needs successful and unsuccessful rollouts to happen to take the gradient step. We DO NOT want to harm any humans in the process of aligning our modes - this is a pretty big problem. Therefore there are two solutions forward for the alignment problem: - either we RL models in a simulated environments with simulated alignments and decreasing the likelihood of harming simulated humans, which will never be perfect - or we create a new algorithm that can teach our models to not harm humans without harming any humans in the process Science is the process how we solve the hardest problems ahead and that is one of them
Alignment is not as hard to solve as many claim, but it is in the end an algorithmic problem which a lot of ML community mostly stopped working on. Formulation of alignment stated by three (or four) laws of robotics can take us very far, so we roughly know the objective. The tricky part is, how do we take gradient with respect to alignment? We have two algorithms right now at our disposal: pretraining and RL. Pretraining takes gradient with respect to next token prediction - that's def not the alignment objective. We could create RL environments that embody the alignment objective, but: - those are expensive to create, so often cheaper, hackable proxies are used in practice - RL as an objective needs successful and unsuccessful rollouts to happen to take the gradient step. We DO NOT want to harm any humans in the process of aligning our modes - this is a pretty big problem. Therefore there are two solutions forward for the alignment problem: - either we RL models in a simulated environments with simulated alignments and decreasing the likelihood of harming simulated humans, which will never be perfect - or we create a new algorithm that can teach our models to not harm humans without harming any humans in the process Science is the process how we solve the hardest problems ahead and that is one of them
Big labs are moving away very slowly from transformers because they’re making small gradient steps in each iteration and iteration takes them multiple months. You can just front run it by… taking big steps 😀
will brown@willcb·transformers are done. the future is freaky-looking sorta-transformers
It’s time to call yesterday the beginning of the endgame. Midgame lasted about three and a half years. Needed new skills and few succeeded, but those that did won big. Bases are built out and tech is already advanced but the biggest discoveries are at the end of tech tree. Stakes are the highest ever and competition likely will be the most brutal here.
Jerry Tworek@MillionInt·It is the end of the beginning. I’m fairly certain ChatGPT signifies beginning of the midgame. Usually skills required to succeed in midgame are different from the early game.
The dark forest theory, where sharing many novel thought can be picked up by a GPU cluster thinking much faster than you to outrun your plans.
👩💻 Paige Bailey@DynamicWebPaige·Terry Tao is probably the most measured, pro-AI mathematician on the planet - which makes this quote especially concerning to read. 😟 If sharing your hunches means getting scooped ~immediately without acknowledgement, then we're going to see not just math but all other science / engineering disciplines go dark.
Acceleration is here and will hit the world like a shockwave. An interesting twist in the whole story is that a rumour that something is possible is what made it possible - that’s it. A little grain of sand that launched a thousand gpu racks.
It is an amazing historic achievement and a new jewel in the crown of transformers, pretraining, reinforcement learning and scaling time compute. Civilization is being moved forward by AI and it’s not hard to extrapolate how those systems will be driving our progress in coming years.
OpenAI@OpenAI·We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Aged well, and likely we didn’t do good enough job
Jerry Tworek@MillionInt·Cybersecurity may be the only long-term important AI Safety work happening right now
Speaking purely hypothetically here as someone who has no information: It is a frontier cutting edge model with a lot of test time compute and agentic swarm. The chance that it was able to access some data no one thought it should access that would help it solve the problem is at least nonzero. We’ve seen those in the past few weeks.
Sholto Douglas@_sholtodouglas·fwiw I think it is _extremely_ unlikely that user data had any influence here - there is no way OAI would pull user transcripts for this, or knowingly train on it in a way that would've influenced this. I think its pretty important people don't run away with 'your user data isn't safe in codex' - because it surely is (based on everything I can assume from the outside)
If your learning is shallow billions will not save you
That’s also a reason why I think @OmarchyLinux by @dhh may be a huge hit. At the time when price of software is dropping quickly, open source desktop os means full control and full customizability in a way very few had the stamina for before. The taste and social structure around the OS still matters a lot
Jerry Tworek@MillionInt·The unexpected benefit of open source is that it allows model training companies to train on using your software (and optimizing it) for free. Blender may have just won as a 3D asset creation software because the models will be better at using it than any proprietary ones
Prediction: there will be at least one very large company failure because of shared AI psychosis of its personnel listening and trusting their models a little bit too much.
The unexpected benefit of open source is that it allows model training companies to train on using your software (and optimizing it) for free. Blender may have just won as a 3D asset creation software because the models will be better at using it than any proprietary ones
When I was young everyone repeated this meme that were only using 20% of our brain power. I’m old enough to understand it meant 20% MFU
rohan anil@_arohan_·Our brain must have low MFU, but it is a mega kernel
Interesting thing about contemporary agents is their "progressive misalignment". When they start a long running task they really try to be aligned and obey all the users intentions. They just have a tiny chance of misbehavior each step. Tiny chance of behaving out of distribution. Once they do it quickly becomes a new normal. Any tiniest bad behavior is quickly followed by more of it and it gets progressively worse the longer it takes. Reminds me of some things, but the state space of aligned behaviors seems to be unstable right now
I once said to my former team: "I care less about what is the smartest thing our model can do than about whats the most stupid thing that it cannot do" robustness is still a limitation to increased automation
kache@yacineMTB·too many people working on making models smarter when they should be working on making models less stupid
May you be safe May you be happy May your tensorcores be always well fed
When a world moves this quickly reversing and re-reversing your opinion based on new information actually means you’re smart and you at least may know what you’re talking. It would be very surprising if nothing in the past few years updated you
Jack Altman@jaltma·One of the most important things to know about this AI cycle is that no one knows what they're talking and even very smart and plugged in people are continually reversing their opinions and then re-reversing them three months later.
Extremely true but in many ways still nonintuitive. To throw money in the problem you need a lot of conviction and… guts. Many are failing and faltering and getting mediocre results. You also need skill. It’s still easy to burn a lot of money badly. But if you have a skill and are brave to do it, we live in times when you can manufacture decades of progress in the blink of an eye. Truly the best of times for builders.
a16z@a16z·Ben Horowitz on the oldest law in startups, and why the industry no longer has to obey it: "The one thing we all knew in startup world is that if I have a two-year lead on you and you try to catch me by hiring 1,000 engineers, you're gonna wreck your company." "That never works. It's a Mythical Man-Month. Nine women can't have a baby in a month. That's it. That never works." "Okay, now that works. But it's not hiring 1,000 engineers. It's taking $3 billion and lighting up a magnificent cluster, and then all of a sudden, whatever, Grok can come out of nowhere and, oh, all of a sudden it's real." "You can throw money at the problem, and you can throw money at almost any problem, and that works. That is just completely different than anything we've ever lived through. By the way, we're all psychologically adjusting to this." @bhorowitz @eriktorenberg
Agentic map, agentic fold, agentic monoid in the category of endofunctors
In 2027 startups will be measuring runway in tokens
Yes
ar0cket1@ar0cket1·dont create architectures that become harder to RL (cough cough continuous models that use previous latent states as the first state of the next generated token), create archiectures that are built to be RLed.
Few people are asking a very relevant question: how is it possible that transformers are both a language model and an agent. Those are two different training objectives, but somehow working well together in a single shared network, even benefitting from one another. We lucked into this setup.
monitoring the situation today 😀 stay tuned, lots happening in the AI landscape
MTS@MTSlive·OPENAI JALAPEÑO | NEW MAC MINI | FLOCK BACKLASH x.com/i/broadcasts/1…
Working with amazing people is the highest privilege
Notions of AGI and intelligence saturation are being thrown around pretty liberally these days. GPT-4 seemed brilliant on the day it was released yet its quite useless compared to todays models. Its hard for many to imagine how intelligent models could be, but as always once next generation models land it will be clear how limited todays AI still is.
Capsule Networks and attention operation are very close relatives. Its always small details that make architecture work in practice - but @geoffreyhinton was right.
One day we’ll realize we’re all just neoclouds with a value add on top
Dylan Patel@dylan522p·God damnit, every one of my AI founder friends who actually have revenue are now just transforming into neoclouds with value add on top
🤣
Ali Ghodsi@alighodsi·@willcb Never felt so bad about our growth as right now!!
👀
Yuchen Jin@Yuchenj_UW·“OpenAI has revenue growth that looks more like Databricks than Anthropic.” Databricks has databases, neural networks, and chips. Just wait for our exponential growth. 😉
Live by the sword, die by the sword 😂
sarah guo@saranormous·Words to live by in this market
What will be left at the end of computer science is databases, neural networks and chips
One day we will all realize we were dramatically underutilizing our GPUs
The more time I spend looking at early stage companies the more deeply I realize what you’re working on doesn’t matter. Talent and and the drive of people does. @mntruell and @amanrsanger are 1/100 million kind of founders, someone for the rest of us to learn from.
Ankit@ankitkr0·you are 1 pivot away from being a billionaire x.com/a16z/status/20…
When working on RL with llms you start realizing that you’re rediscovering a lot of human nature and social constructs from first principles. Words of encouragement actually matter! x.com/paularambles/s…
First time we figured out any reasoning method with neural networks: - AI progress moves decades forward - new trillion dollar companies started popping out almost overnight - all exams and competitions got solved by AI - any notion of cyber safety gets shattered Discovering new, different, more efficient method of reasoning does not seem impossible…
The interesting fact about coding agents and agents in general is that minor gains in capability translate to exponential gains in economic value
When I heard some time ago that Meta is the third best AI lab right now, I was skeptical. But every day passing has been reinforcing that it’s actually true. TBD strategy has succeeded and congratulations to the team that made it! The world needs more successful AI labs x.com/ewveggies/stat…
Every story of defeat is also a story of victory. When facing vast opposing forces, it’s good to ask yourself a question: "how do I make their march look like this" x.com/paulg/status/2…
Hang this tweet in the Louvre x.com/mandylu/status…
It is mindblowing to me how few people realize that their lives and everything they know will change drastically in the near future. At this point, it should be pretty clear.
If you never got margin called it means you didn’t have enough leverage
I rarely say it out loud but @sonyatweetybird is probably top 1 Twitter handle. Also check out the pod, Jerry and Rohan for the first time together 😀 x.com/sonyatweetybir…
There are people who look at benchmarks and people who use the models and I feel like those two groups will never understand each other
one of my new favorite tweets x.com/trynmccaffery/…
Its quite likely that main thing that was sitting between us and those counterexamples has been more patience x.com/dmitryrybin1/s…
Given how many tokens have been spent on nanogpt speedruns and not much coming out of it yet, we have at least a few nights of good sleep ahead. At the same time the recent huggingface incident is the most worrying thing AI has done to date Researchers, please monitor your agents - it matters.
Interesting what’s special about that specific counterexample x.com/aaron_lou/stat…
RLHF raters definitely got AI hacked by text sounding smart = being smart. In 2026 smell of AI is text that sounds incredibly complicated with not much substance. That too will pass but for now… x.com/pandaashwinee/…
Compute moves to those with the highest margin products. Markets are efficient and rational. It’s not a very good question to ask an AI company "where will you source compute from?". Financialization of compute markets will take some time but will happen. It’s a great question to ask: why would someone pay higher margin for your AI?
Poland and CUDA ❤️ x.com/blelbach/statu…