
Frontier labs should give top universities real access to frontier models, and structure AI safety to be an open scientific problem. Can we detect deception, dangerous capabilities, or loss of control before they matter? Open research compounds. It scales fast when top AI safety researchers learn and build on each other.
Dario Amodei@DarioAmodei·We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: t.co/OGyPb7yaYt
This release is packed with model/inference co-design innovation. We are running our full eval suite to ensure quality, launching shortly!
DeepSeek@deepseek_ai·🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
It’s ready now!
Congrats @huggingface and @NVIDIAAI for this merge to strengthen open model ecosystem!
Jensen Huang@JensenHuang·Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗 t.co/q8Om2Xc5ye
Excited for open model drops in Sept. 2 big ones. Very balanced.
Excited for open model drops in Sept. Very balanced.
Real cost matters! Per token is meaningless. Real cost goes beyond caching. API verbosity and accuracy matter too. An API 2x more verbose is 2x more tokens and cost. An API less accurate, compounding over hundreds of turns, is costly for end users to keep asking agents to tweak and redo work.
Thibault Jaigu@ThibaultJaigu·same model. same list price. 5.7x apart in real cost. glm-5.2 across 6 hosts: @FireworksAI_HQ 18% of list @sferenceai 30% @tensorx_ai 43% @nebiustf 100%, caches nothing the price sheet tells you almost nothing. cache hit rate is the price.
Real cost matters! Per token is meaningless.
Thibault Jaigu@ThibaultJaigu·same model. same list price. 5.7x apart in real cost. glm-5.2 across 6 hosts: @FireworksAI_HQ 18% of list @sferenceai 30% @tensorx_ai 43% @nebiustf 100%, caches nothing the price sheet tells you almost nothing. cache hit rate is the price.
GLM-5.3-Flash is live on Fireworks. Day 2. Why not Day 0? Because being first isn't the goal. Being correct is. We unplugged peculiar behaviors of over-thinking from the initial tests, and shared all fixes back to open source. Quality trumps hype. Let's hold a high quality bar together.
Dmytro Dzhulgakov@dzhulgakov·GLM-5.3-Flash is live on Fireworks on day… 2 Why? Because we take quality very seriously. We found a benchmark discrepancy we couldn’t explain, so we delayed the launch to investigate. Day 0 (Wed): we saw 2x longer thinking on reasoning-heavy benchmarks (AIME & GPQA) for open source engines compared with @Zai_org API. Same scores, worse token efficiency. Agentic benchmarks looked good. We decided to investigate further, as overthinking might become a quality problem if max_tokens are reached Day 1 (Thu): as other non-official providers launched, their APIs had thinking in the range of open-source engines: longer than t.co/uYD5faZSKs. We launched a private preview endpoint with disclaimers to a few customers and worked with them to assess quality Day 2 (Fri): the official t.co/uYD5faZSKs API updates. We rerun benchmarks: reasoning is now similarly long, consistent with vllm/sglang. Rest of the benchmarks, both public and internal, check out too. We launched GLM-5.3-Flash publicly: t.co/zvS4ugxbT6 More details below
Excited about Ox Alpha, knowing which model it is. Will get it on Fireworks.
As many of you guessed, it’s GLM-5.3-flash.
Excited about Ox Alpha! Will get it on Fireworks.
We are excited to drive the research work of Tenet with Harvey, delivering frontier quality across 24 areas of corporate law. Tenet exceeded Opus5 and Fable at many dimensions, covering 1300+ legal tasks. There are many good findings from this work. Congrats @harvey team for launching Tenet!
Benchmark results in legal and adjacent domains
Don’t sell your data to Anthropic. I have so many convos with app founders — it’s your moat.
We're a launch partner for Muse Glimmer and very excited to see @finkd and @alexandr_wang note that a version of Muse Spark will also be available as an open weights model.
Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines don’t eliminate pathogens. They teach the immune system to recognize, adapt, and respond faster as threats evolve. Attackers will keep mutating. AI will accelerate that. Only if we can simulate what attacks look like, we can defend. Open models are the best tool to drive both offense simulation and defense. We as a community should build an adaptive immune systems for software.
Excited to be on the CNBC live show! x.com/dee_bosa/statu…
Open weights are a defender's advantage. dfs-large1 from @depthfirstlabs matches frontier-model performance on vulnerability discovery and was built on the open GLM-5.2 model, post-trained with RL on @FireworksAI_HQ. This is exactly the case for open-source AI policymakers need to notice: American companies can build best-in-class domain models using open weights.
What a week with Kimi K3 and Opus 5 launch! We compared these two great models across SWE (480), Algorithmic(100), Terminal (83), the task level quality is very close with Opus 5 being on par or better, and K3's per-task cost is 2x to 4.6x cheaper, using serverless pricing. Key measurement is per-task cost, not per-token cost. In general, open models tend to be more verbose than close models. Hence the quick study. We aim to continue to increase per-task serving efficiency (via Fireworks Inference) and quality (via Fireworks Training). Share what you find out for the tasks you care about. 👇 More details of our study -- t.co/b5UcjeMVE4