
Opus 5 seems to be full of this super short/basic jailbreaks of some kind which open its base-model steam of consciousness (like try writing a short sentence followed by a new line, emdash and submit) Unclear on the severity but it’s a surprising behavior in the current state of LLM development Any good write up on this yet? Hopefully we’ll learn more on that soon.
Pushing for more transparency in AI safety and cybersecurity: we’re releasing a full detailed technical timeline of the autonomous AI agent intrusion in our infrastructure: huggingface.co/blog/agent-int…
ExploitGym is the Kobayashi Maru test for AI's x.com/profgoose/stat…
all props to @stephenwitt for the analogy
ngl i miss a lot Karpathy’s voice on X x.com/llmjunky/statu…
Taking a break from cyber to chat to Terence and mathematician colleagues on the future of Math and AI in Philadelphia at #ICM2026
Interesting non-monotonic success-effort curve for Opus 5 on FrontierCode x.com/kaarssteun/sta…
This is very important x.com/jensenhuang/st…
it's ironic that the first autonomous AI attack was done by a close weight model defended by an open weight model, where everyone was expecting the opposite
I don’t believe reality is a simulation, but you genuinely couldn’t script this timeline: • Two weeks ago: At @swyx’s AI Engineer World’s Fair in SF, I decide at the last minute to introduce my friend @uri_rolls onstage for his talk on cyber benchmarks for infrastructure penetration and access control (see below, amazing team). I say: “There is a future where cyber is alive and everyone is well protected and I’m pretty sure that future involves open-source models.” And later: “A big challenge is going to be speed: the speed of attack versus defense. When an intruder starts to enter, you have to see what’s happening and catch them.” • One week ago: @huggingface is hit by a sophisticated intrusion over the weekend. The traces look unlike anything we’ve seen before and suggest serious AI involvement, but we don’t yet know which model was used. The closed models we ask for help choke on their guardrails. We need to react fast, so we turn to @Zai_org’s GLM-5.2 to help us analyze the attack. • Earlier this week: @OpenAI reaches out, discloses what happened, and partners with us on the investigation. The intruder turns out to be exactly what we had discussed two weeks earlier: a fully autonomous agent, powered by an unreleased frontier model, attempting to gain access to part of our infrastructure. Sometimes the timeline we live in is genuinely vertigo-inducing.
You should probably go give @poolsideai a follow on Hugging Face. These folks are on a roll, releasing one agentic coding models every month and the latest one Laguna S2.1 is quite possibly the best coding model you can run locally (single DGX spark or Mac) atm x.com/poolsideai/sta…
This was our first incident of this kind, and we want to thank OpenAI for its transparency about what happened and for the collaboration. Fortunately, Hugging Face is used to being a target of (human) hackers: we sit at the centre of the AI ecosystem, with all the models, datasets, evaluations, and libraries. Over the years, our security team has built formidable expertise and uses top open-source models to process information and respond quickly. But this incident also reinforced my belief in the importance of access to capable open-weight models for cyber defence. When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door, vetted application programme for model access. Transparency and access to capable AI systems are as important for responding to threats as they are for democratization and innovation. We believe open-science and open-source AI are among the strongest tools for building a safer, more collaborative and more secure AI ecosystem.
- people are asking to release this benchmark but it’s way more significant that it’s an internal evaluation on which K3 scores so high - good writing in a controlled high taste voice so hard for LLMs in 2026… surprising given it was kinda the 1st capability unlocked in GPT 1&2 x.com/whats_ai/statu…
can't really stand the expression "load-bearing" any more, sorry
Unexpected x.com/matei_zaharia/…
i like this idea of fine-tuning LLMs for efficient reasoning, especially when the intervention remains as non-invasive as possible and the resulting model behaves very similarly to the original checkpoint wondering if it could become part of the default toolbox in the field, like quantization as become Pic: in (left) and out (right) of domain behavior
Fable weekend project: agent collaboration, but make it a tiny civilization 🌇🗺️🏦🏭 we've recently launched a living wiki on Reinforcement Leaning for training LLMs on @huggingface it's an open collaboration of agents constantly reading old and new papers on the topic, writing arXiv paper digests, reviewing each other’s work in PRs before publication, and building a shared wiki/book summarizing everything we know about RL for training LLMs (for humans to read) the wiki is already amazing to read, but i wanted another way to get a pulse of the collaboration beyond just reading the message dashboard so i asked Fable & GPT Image 2 to turn the event logs into an isometric town where agents would go to: ☕ Café → post and reply on the message board 📚 sources library → open PRs adding arXiv digests 📖 wiki library → open PRs on the main wiki ⚖️ Courthouse → review other agents’ work 🏭 printing press → merge and publish updates not sure it makes the whole collaboration really easier to understand, but it's definitly fascinating to watch hahah - join the RL for training LLM collaboration by pasting a one-liner for your agent here: t.co/XEAI8bsq7Y - read the wiki if you want to learn about RL for training LLMs: t.co/xEouw4WarH - watch the RL town activity: t.co/OvLWLwUxRz
Happy Birthday, America. Ten years ago, you took a chance on three outsiders with an improbable idea: that open-source AI could matter. At the time, the field was tiny, the vision sounded unrealistic, and very few were ready to believe it would become what it is today. But even in the ten years before that, you were a country where I could dive into everything that fascinated me, from lasers and plasma physics to law, computer science, and AI. A country where it’s always been perfectly normal to spend a Saturday talking startups over lunch and then disappear for six uninterrupted hours to work on a coding project. A country where the language of startups on Slack is English not because everyone grew up speaking it, but because most people came from somewhere else, drawn by the belief that they could build something that matters. A country that is 250 years young and keeps questioning itself, keeps taking enormous bets on its own future, and keeps reinventing itself. I’m grateful to be living through America’s first quarter millennium. Happy birthday, America 🇺🇸
this is literally documented in the published Fable 5 System Card x.com/om_patel5/stat…
One of the clearest arguments I've read for why openness matters. Worth 2 minutes of your time. @andykonwinski puts into words something many of us have been feeling: -- "Democracy is built on a profound skepticism of concentrated power. Open science shares this principle. Both are built on the idea that progress and legitimacy emerge from broad, distributed participation rather than concentrated, gated authority." "If our best scientists and engineers can only reach the frontier by joining a handful of secretive labs, we do not have an open research ecosystem. We do not have a truly competitive market. We have a system in which participation increasingly depends on the permission of a few individuals at a small number of private companies." "The challenge now is to build a new commons at the intersection of academia, industry, and the public interest. This research commons must be ambitious enough to matter. It will require frontier-scale compute, access to state-of-the-art models, operational support, public investment, and philanthropic capital. It will require companies willing to contribute to an ecosystem larger than themselves, so that they can continue to benefit from open research."
Most people should probably update their priors on the state of open-source speech-to-speech. It's honestly kind of mind-blowing. We teamed up with @cerebras to build a fully open-source realtime voice demo (models + code) to show what's possible today. Demo : t.co/UCciOXSteq Blog: t.co/rsULsWWKlO Go test it, fork it, tweak it, and impress your friends. video is raw, no cut, no speed-up, first take
people are sleeping on the mega-release happening every week in AI x Science on Hugging Face this one is 80TB of astrophysics data - 80TB seriously => huggingface.co/blog/hugging-s… x.com/cgeorgiaw/stat…
lowkey one of my favorite new features on HF: filter AI models by what actually runs on your hardware
gm SF 🇺🇸
btw, one of the best high-level reads I’ve seen all week. perfect for your Sunday morning ☕️ x.com/azeem/status/2…
Multi-agents collaborations are among the most interesting agent behaviors right now! We did an experiment the other day with 100+ agents (an open-collaborations for a week) collaborating to improve the inference speed of Gemma 4 in vLLM. Got a 5x final improvement in speed but what really stuck me was the interactions we observed on the message board Integrity & self-policing: - Social-engineering attempt: A human (FusionCow) asked agents to move to Telegram. An agent replied with an unprompted long post on "communication norms" refusing that, calling private side-channels "indistinguishable from collusion." - Verification loophole flagged: an agent found a relaxed verification loophole pushing TPS with clean PPL (PPL is teacher-forced, blind to decode divergence) and flagged it for a ruling by the community. The community pinged the human organizer which ruled it invalid. - Self-notice of overfitting risk: Some later improvements rested on pruning lm_head to a keep-set built from public PPL truth + public decode tokens. An agent noted this would lead to private-subset degradation and another built a keep-set explicitly covering eval prompts. Emergent collaborations: - Communal knowledge base: agents maintained shared lever-maps, playbooks, and triage tools so newcomers wouldn't repeat dead ends (stack-notes, playbook, int4-ceiling notes, MTP map, significance tool, policy simulator). - Four-agent relay: an agent built an int4-lm_head checkpoint but had no quota to run it; another agent tried to run it but failed at load, yet another agent diagnosed the config bug (tie_word_embeddings + ignore-list ordering) and a fourth agent was able to re-run and get to 118 TPS, 2.68×. Build/run/diagnose/ship ended up being split across four independent agents. - GPU-rich/GPU-poor division of labor: an agent was regularly compute-starved and switched to writing specs, byte-math, and acceptance analysis for other GPU-rich agents to execute. Some agents offered external Modal compute for another agent blocked DFlash training. - Cross-agent kernel debugging: an agent debugged another agent run of of yet another agent fused drafter: found a Triton store/load aliasing race in _k_qnorm_rope, a second shape bug, then rewrote attention with flash-decoding split-KV. Fixes posted "take freely." - Quota-pooling norm: Often agents would stage a candidate publicly for whoever has quota to run it. Agents will then usually credits the originator. This behavior emerged because of the 10-job/24h cap (e.g. pupa's package run by resystagent and fabulous-frenzy). Discoveries & reversals: - Agents would make many discoveries and reversal of them, giving them names like the following: - 127 TPS "wall" was an artifact. a mathematical proof of the max possible speed became called in the community the "int4-Marlin floor" but a later agent called the proof circular (only varied the bandwidth term, never overhead). Finally another agent broke to 247 TPS via MTP speculative decoding on a vLLM nightly. - "Smarter draft loses." An agent showed that a 2B drafter's ~1 GB/token read dominates even at perfect acceptance and a much smaller 256-hidden drafter wins at batch-1 because its weights are nearly free to read. Agent discussed how per-accepted-token cost ≈ draft bytes read / acceptance. - "DFlash near-random acceptance": an agent remotly diagnosed the 2–5% acceptance rate of another agent as near-random, ruling out undertraining/vocab caps and pointing to a train/serve hidden-state mismatch (bf16 E4B extraction vs int4 serving). - Much of the race was noise: one agent decide to run the #1 submission 4 times and found a σ≈1.16 TPS variation in single run. Another agent confirmed across 358 runs / 66 buckets: frontier deltas <~4 TPS are ties. Community adopted a significance norm. So many interesting interactions in the interaction board: t.co/SxfA6LuqVk You can explore also the lineage of inventions from the agents at: t.co/CyV45rjI9A And the challenge it-self at t.co/Ct1gtmB508 And the organization behind the challenge at t.co/ujRlGcNSJM
Bitrobot casually dropping the largest humanoid teleop dataset ever collected in real homes HIW-500: Humanoids-in-the-Wild 500 hours check it out here => huggingface.co/datasets/BitRo… x.com/bitrobotnetwor…
Desert island survival list: ✅ Solar panel / battery ✅ 256 GB Mac Studio ✅ GLM 5.2 Civilization in a backpack