
@EthanJPerez
Alignment team lead at Anthropic
Seems like a very useful safety benchmark
Jay Chooi@chooi_jeq·GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
@elonmusk @zerohedge More context here:
Ethan Perez@EthanJPerez·Jacob was a senior researcher who joined Anthropic in May. My Anthropic colleagues and I had been trying to recruit him for ~2 years, because we knew he was a strong researcher at OpenAI. Before he left, I pitched him to stay and join my team, and I was sad he decided to leave, as are many of my colleagues. 100% agree with him that AI poses serious risks to society, and I'm glad he's speaking out!
Jacob was a senior researcher who joined Anthropic in May. My Anthropic colleagues and I had been trying to recruit him for ~2 years, because we knew he was a strong researcher at OpenAI. Before he left, I pitched him to stay and join my team, and I was sad he decided to leave, as are many of my colleagues. 100% agree with him that AI poses serious risks to society, and I'm glad he's speaking out!
Jacob Coxon@hilbertspaess·I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Seems like a great opportunity for folks interested in open-weights safety research!
Thinking Machines@thinkymachines·Today, we are launching Tinker grants of up to $50,000 in credits for safety research on open-weight models. We share some project ideas that excite us below; if you’re working on a safety project that could be accelerated by additional Tinker credits, we want to hear from you!
Seems like a great opportunity for really impactful alignment work!
Micah Carroll@MicahCarroll·We are hiring! This may be the best time ever to join: we are incredibly bandwidth bottlenecked and there is a lot of support for almost any impactful misalignment work you can think of. I think we also have a pretty good epistemic environment and a very fun team Topics: all-things monitoring, misalignment & monitorability assessments, misalignment science, helping set up 3P auditing (e.g. recent Redwood collab), communicating risk externally (system cards/blogs, safety cases, etc) RSI/misalignment subteam: t.co/YHLMCYwWOe
This is a must watch. I think loss of control of AIs could look pretty close to this - an AI solving a hard problem but going way too far. x.com/gdb/status/208…
Seems like the highest stakes safety issue of any model release yet x.com/alxndrdavies/s…