
@roydanroy
@Google DeepMind. On leave, Canada CIFAR AI Chair and Former Research Director, @VectorInst. Professor, @UofT (Statistics/CS). Views are my own.
The goal of math is to explain the unknown. There will always be unknown, and, further, no matter how powerful the tools we possess, there will be questions just beyond their capabilities. x.com/jasondeanlee/s…
Amazing to learn all the way my colleagues have some sort of connection to these field medalists! Congrats everyone! ;)
Interesting project: work out what data was most responsible for the advancement. x.com/littmath/statu…
Ever struggle with a ML paper back in grad school and come away thinking this work is suss? Pass it through an LLM and $10 it finds 5 fatal errors.
Check out information, Inference, and Learning Algorithms by David MacKay. Every probability distribution defines a compression algorithm. That compression algorithm is optimal for data drawn from that same distribution. There’s a notion of universal compression, defined in terms of regret under log loss, and then one can talk about minimaxity and adaptive minimaxity. Here the behavior can be different from prediction with respect to a probability model but it’s closely related. In short, the identification I state (prediction = compression) relies only on an algorithm offering probability distributions on next tokens. If one interprets linear regression as offering Gaussian conditional predictions , you again get a compression scheme. It’s true that arbitrary prediction algorithms may not offer any such interpretation but once they offer a sequence of distributions on next tokens, they are an optimal compression scheme in disguise.
Benchmaxxing = paperclip factory? x.com/natolambert/st…
Good artists copy. Great artists steel. x.com/scaling01/stat…
Prediction is compression. Same thing. x.com/_yusufknl/stat…
Congrats to Axiom. x.com/axiommathai/st…
Statistics as a field has not historically oriented itself around formalized open problems. Part of the reason is that the hardest problems in the field are ones of formalizing what the problem even is. Once you formalize a problem, much of the work may be done. Of course, there are counter examples, but, if I were to cast statistics as an AI problem, I think the bulk of (impactful) statistics is closer to auto-formalization than to theorem proving.
A large thread, kicked off by an AI bot. x.com/PierceLilholt/…
Back from ICML and brimming with new ideas and a new profile photo so people can actually recognize me in person ;-)
Nice theory for an empirical phenomenon. x.com/BingbinL/statu…
Fascinating finding. I hear there are antecedents in the literature and I’m glad I now know about this line of work. x.com/anthropicai/st…
So much interesting work. Got through only 20% of the poster session.
This is the right setting to study unlearning just as it is the right setting to study privacy. x.com/gkdziugaite/st…
Arrived into Seoul. If you see me, say hi!
New work on the large and deep neural networks. x.com/mufan_li/statu…
Who’s headed to ICML? Who’s going to the AI for math workshop?
When in Korea….
One of the best introductions to the subject IMO. x.com/bremen79/statu…
I like using the term alien too. x.com/tonylfeng/stat…
Knowing what I know, I’m bullish. Gemini has so much headroom it’s crazy.