
@jacobandreas
Teaching computers to read. Assoc. prof @MITEECS / @MIT_CSAIL / @NLP_MIT (he/him). https://t.co/5kCnXHjtlY https://t.co/2A3qF5vdJw
👉 New preprint! (and a direction I'm excited to do more work on soon) x.com/elinorpd_/stat…
@kanishkamisra @evanqed @arnab_api I am now taking credit for everything this guy did. (But seriously the advising credit for the paper goes to @davidbau @boknilev!)
If you're excited about Anthropic's J-space work, definitely worth checking out the original paper on Jacobian lenses by @evanqed and @arnab_api! arxiv.org/abs/2308.09124
👉 New preprint (we had a big backlog 😅)! Revisiting adversarial imitation learning for the era of RLVR: x.com/MehulDamani2/s…
👉 Preprint: understanding learning dynamics & mechanisms in LMs trained to explain / predict their own behaviors! x.com/CarlGuo866/sta…
👉 New preprint! Automated interpretability by approximating / replacing NN components (here attention heads) with programs. x.com/amirihayes_/st…
👉 New preprint! Optimizing LMs so that *how they describe themselves* matches *how they behave* (with applications to alignment, explainability, and building structured models of complex data generating processes) x.com/PresItamar/sta…
👉 New preprint on revisiting continuous-space diffusion models! x.com/linluqiu/statu…
@jiaxinwen22 @PresItamar This actually reminds me of experiments @StephenLCasper did back in the day on doing full-fine tuning with a CCS-style loss, which I suppose would be the direct mapping onto the proposed objective here. Results were mixed but not sure if anyone has revisited this with modern LMs.
👉 New preprint: how do we make LMs more reliable once there's no more training data? Enforcing *consistency* of LM predictions across inputs lets us unsupervisedly optimize for factual accuracy & faithful explanation (& get a unifying view on many existing post-training algs) x.com/PresItamar/sta…
Sad to be missing #NeurIPS2025, but check out: - @MorrisYau's oral on linear transformer learning w guarantees neurips.cc/virtual/2025/o… - Christy Li's paper on red-teaming agents neurips.cc/virtual/2025/l… - @ReeceShuttle's paper on what LoRA learns neurips.cc/virtual/2025/l…
2nd paper is @christy_li_!