
@FireworksAI_HQ
The frontier platform for training and inference on open-weights models at scale.
Today we're launching the Specialized Intelligence Index (SII): one destination for real-work benchmarks across industries, built by the teams that use them every day. Hear from Fireworks co-founder @the_bunny_chen on the importance of specialized benchmarks:
Featured on SII: @doximity @mercor @harvey, @depthfirstlabs, @novee_security, @genspark_ai, @SierraPlatform, @DecagonAI, @ProximalHQ, @Macroscope, @RogoAI, @traversal_ai fireworks.ai/index
We ran 18 models across 113 real coding tasks on DeepSWE, then went back and asked a simple question: what if every task had routed to the model that handled it best? Answer: 97.6% solve rate at $1.88 per task, versus the best model at 74.1% at $6.52. The next frontier is a router. Full analysis: t.co/yzpA9I3V39
Don’t miss newly announced Fireworks Forge speaker: @BrendanFoody, CEO and Co-founder of @Mercor. Join the people, teams, and companies building their own frontier on open models. Nov 3, San Francisco. Apply to attend: fireworks.ai/forge#apply
Behind every production agent is a serious stack. Join @e2b, Fireworks, and @braintrust on 9/30 for technical talks on sandboxing untrusted code, serving models at agent-loop speed, and knowing whether a change actually helped. RSVP: luma.com/e2b-0e34
We’ve just added three new speakers to our incredible Fireworks Forge lineup! We are proud to welcome @BrendanFoody @pirroh and @jefftangney to the stage as we bring together the people, teams, and companies building their own frontier on open models.
Apply to attend: fireworks.ai/forge#apply
Open models are becoming increasingly capable in biology. Serving them well is harder. Fireworks gives @phylo_bio day-zero access and a fast serverless path, putting research agents in front of more scientists. Learn more: fireworks.ai/blog/phylo-bri…
Come hang out with AI engineers, tech leads, and the Fireworks team during AI Engineer 2026. Want to leave with hands-on experience fine-tuning, evaluating, and deploying a model on Fireworks? Bring your laptop. We'll make it happen. Join us: luma.com/gydwrqt6
GLM 5.3 is now available for training on the Serverless Training API, joining Kimi K3, Qwen 3.8 27b and more. GLM 5.3 Flash coming soon. No sales demo or DMs required. Here to help as needed, but you can also spend that time building. Start with docs or a pre-made recipe ➔ t.co/URmbsVlNp2
Quickstart docs: docs.fireworks.ai/fine-tuning/tr…
Many assume that agent spend goes toward output tokens. When we ran DeepSWE on Astra vs. DeepSeek V4.1-Flash, input tokens outnumbered output 174 to 1. 99.6% were cache hits. Those hits are 60% of the bill. Net result? Same quality. $0.43/task vs $6.52. x.com/i/article/2099…
Speed is key for @juicebox_work. They need to search hundreds of thousands of talent profiles in seconds with multiple real-time search agents. We helped them cut latency 80% and drop inference costs from $4M to $800K/yr, using specialized models tailored to their use case.
Excited to continue our partnership with the @alibaba_cloud team on this one, bringing the latest version of Qwen's flaship model to the Fireworks community. Start building with Qwen3p8 today: fireworks.ai/models/firewor…
Alibaba Cloud@alibaba_cloud·Developers need speed and uncompromising production quality. 🚀 We’ve partnered with @FireworksAI_HQ to deliver exactly that for Qwen 3.8 Max. Build faster, scale effortlessly. 🤝 #Qwen #FireworksAI #LLM #OpenSource
Reinforcement learning at @cognition's scale is a hard infrastructure problem. We are proud to be part of the stack behind it. Congrats to the team on SWE-2! Read more about how we think about RL at Fireworks: fireworks.ai/blog/frontier-…
Cognition@cognition·Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
Recently, we hosted our 12th Nerd Meetup, this time at @lightfld 's office in SF. Builders came out from Lightfield, @qdrant_engine, and Fireworks. Building in the AI/infra space and want to join us for the next meetup? Follow along. We'll announce the next one soon!
GLM-5.3 is now available for training on Fireworks Dedicated Training API and Managed Training surfaces. Built for complex coding and long-horizon agents and now available to train on our dedicated infrastructure. Get started today: docs.fireworks.ai/fine-tuning/tr…
GLM-5.3 is now available for training on Fireworks' Dedicated Training API. Built for complex coding and long-horizon agents and now available to train on our dedicated infrastructure. Get started today: docs.fireworks.ai/fine-tuning/tr…
Deepseek-V4.1-Flash is available now on Fireworks! It is a 552B Parameter MoE built for coding, cybersecurity, and agents. It is the ideal workhorse model that outperforms Opus 5 and GPT 5.6 Sol at 1/40th the cost on DeepSWE,CyberGym, and Automation Bench! Interested in higher quality? Reach out, as we’re bringing this model to Fireworks Training soon. Try Deepseek-V4.1-Flash now on Fireworks: t.co/2Bez044BvU
Deepseek-V4.1-Flash is available now on Fireworks! It is a 552B Parameter MoE built for coding, cybersecurity, and agents. It is the ideal workhorse model that outperforms Opus 5 and GPT 5.6 Sol at 1/40th the cost on DeepSWE,CyberGym, and Automation Bench! Interested in higher quality? Reach out, as we’re bringing this model to Fireworks Training soon. Try Deepseek-V4.1-Flash now on Fireworks: t.co/2Bez044BvU
Introducing Gen-1 Slides: an open model that matches Claude Opus 5 on slide generation, at ~1/17 of its input-token price. @genspark_ai post-trained it from a @MiniMax_AI M3 base with Fireworks Lab using long-horizon RL. Live today as the default in Genspark AI Slides. 🧵
Genspark owned the product judgment. They brought their real production environment into training and shaped what the model optimized for: the standard for a good deck, the design principles behind it, and how to sharpen that standard over time.
Today, 10am PDT: we'll cover the journey from renting closed frontier models to owning your specialized intelligence, with a practical framework for knowing when to make the leap. Includes open Q&A w/ Head of AI Developer Education @Prof_OZ Register: fireworks.ai/event/webinar-…
Renting a closed model + a brittle, overfit harness only gets you so far. Routing to open models helps quality, cost, and latency, but you hit a ceiling. Most ambitious AI teams train their taxonomy, style, and judgment into models they own. Read more: fireworks.ai/blog/making-th…
We invite you to Forge. At Forge, we will challenge you to make your own frontier. You'll spend the day with the builders, researchers, and technical leaders pushing AI forward. Lin Qiao. Jensen Huang. Jay Parikh. And more. Apply to attend: fireworks.ai/forge
We invite you to Forge. At Forge, we will challenge you to make your own frontier. You'll spend the day with the builders, researchers, and technical leaders pushing AI forward. Jensen Huang. Lin Qiao. Michele Catasta. And more. Apply to attend: fireworks.ai/forge
Congrats to one of our Kimi K3 Fire Pass Community Hackathon winners @sethsaler for his winning entry: Dayprint! Our Head of AI Developer Education, the esteemed @Prof_OZ, walks through Seth's entry with live commentary. #kimik3onfireworks
Seth Saler@sethsaler·Inspired by all the cool projects I've seen lately on X, dayprint brings your personal context together to plan your day, review your performance, and synthesize how your week went. My entry for #KimiK3onFireworks ↓ @FireworksAI_HQ x.com/FireworksAI_HQ…
Our partners at @CosineAI train their own coding model on real production code, for languages like Fortran and Verilog that generic models fumble. On cost per successful task, their Lumen Outpost model beats GPT-5.5 by more than 3x. The whole pipeline runs on Fireworks Training. Build your own frontier: t.co/E2RBk9thSV
Our Head of AI Developer Education @Prof_OZ is back with a new series! He’ll cover the journey from renting closed frontier models to owning your specialized intelligence, with a practical framework for knowing when to make the leap. Register here: fireworks.ai/event/webinar-…
Our co-founder @the_bunny_chen sat down with @FirasSozan to talk about why strong open models increase compute demand, not the other way around. Watch the episode: youtu.be/qb58MI9i1nE?si…
Research doesn’t end in right or wrong, but in insights that lead to more questions. General-purpose models aren’t fit for research because they don’t know customers’ workflows. @quadrillionlabs trains large custom models on its proprietary data using the Training API so its agent fits how researchers actually work. Build your own frontier: t.co/E2RBk9sK3n
We're going to need a bigger park. Fireworks is hiring across all teams. Check out our careers page, and come shape the future of AI with us! See you soon: fireworks.ai/careers
A noisy code reviewer gets ignored. Flag too many non-issues and engineers tune it out. @Macroscope got the control they needed with the Fireworks Training API to train a model using their own advantages and loss to improve the signal-to-noise ratio. Build your own frontier: t.co/E2RBk9thSV
Our partners @LangChain generate billions of tokens of agent traces a day, their richest signal on how real users react. Judging every one with a frontier closed model like GPT-5.5 or Opus is too costly, so they fine-tuned a Qwen base model on Fireworks that matches it at up to 100x lower cost. Build your own frontier: t.co/E2RBk9sK3n
DeepSeek V4 Flash now comes with native vision on Fireworks. Available now on serverless and priority tiers. Better benchmarks than the text-only 0731 version at the same low price: $0.22 input / $0.66 output per 1M. Give your workhorse model eyes. Try it today → t.co/ubE2Z2AnZ5
DeepSeek V4 Flash now comes with native vision on Fireworks. Available now on serverless and priority tiers. Better benchmarks than the text-only 0731 version at the same low price: $0.22 input / $0.66 output per 1M. Give your workhorse model eyes. Try it today → t.co/ubE2Z2AnZ5
Train past the frontier. @tryheidi's ambient scribe drafts a clinical note the clinician has to trust enough to sign. They fine-tuned open models on Fireworks whose notes their clinicians preferred over Gemini’s in side-by-side reviews, at 3.5x lower latency. Build your own frontier: t.co/E2RBk9thSV
A secret that leaks into shipped code is a breach waiting to happen. @FactoryAI's droids write and commit code faster than any reviewer can keep up, so Droid Shield has to be trustworthy and low-friction: catch the real secrets, clear the false alarms. With the Training API, Factory fine-tuned an open Qwen model that caught almost 20% more real secrets than GPT-5.5 at lower cost and latency. Build your own frontier: t.co/E2RBk9thSV
All week we’re spotlighting teams training past the frontier on Fireworks. @Figma’s AI team is shipping features that change how millions of people design, and with the Training API they can iterate faster. Less time debugging infra, more time on research. Build your own frontier: t.co/E2RBk9thSV
The most ambitious companies are training models to outperform the frontier on the capabilities that differentiate their business. Today, we announce the general availability of our Training API and Fireworks Lab, making model specialization accessible to all organizations.
Working with ML teams around the world surfaced where training workflows break down: constrained model and method choice, limited loop control, idle and costly compute, and split training and serving infrastructure that destabilizes runs Those partnerships shaped our Training API.
GLM-5.3-Flash: GA on Fireworks. We held it two days over a reasoning-length gap we couldn't explain. Same weights ≠ same model. Quality first. Start building → fireworks.ai/models/firewor…
GLM 5.3 is now live on Fireworks, day zero: → 50% improvement over GLM 5.2 on t.co/wQkYTEKzIU’s Code Bench → State-of-the-art open-weight performance on cybersecurity tasks → Excels at complex coding and long-horizon tasks → US-Hosted Serverless on Fireworks Try it now: t.co/9RTwsylMMA
GLM 5.3 is now live on Fireworks, day zero → 50% improvement over GLM 5.2 on t.co/wQkYTEKzIU’s Code Bench → State-of-the-art open-weight performance on cybersecurity tasks → Excels at complex coding and long-horizon tasks → US-Hosted Serverless on Fireworks Try it now: t.co/9RTwsylMMA
Congratulations to @LotusDecoder for their winning submission to the Kimi K3 Community Hackathon. Their entry, MedRecap, gets the full @Prof_OZ hackathon judging treatment here:
Sinan Ozdemir@Prof_OZ·Congratulations on being one of our winners! I took a look and was very impressed :) I'm curious where you got the inspiration for this! x.com/LotusDecoder/s…
Recently @Harvey introduced Tenet, its first model, trained for long horizon legal work. Harvey post-trained it from a Kimi K3 base in collaboration with Fireworks using async RL on our Training API. A thread on promising initial results for both performance and cost-efficiency🧵
On LAB, Harvey's Legal Agent Benchmark, agents produce finished legal deliverables graded against dozens of criteria. All-pass counts a task only if it clears them all. Tenet lifts all-pass from 10.8% to 19.7% over the Kimi K3 base, reaching state-of-the-art on LAB Contracts.
Recently @Harvey introduced Tenet, its first model, trained for long horizon legal work. Harvey post-trained it from a Kimi K3 base in collaboration with Fireworks using async RL on our Training API. Promising initial results for both performance and cost-efficiency:🧵
On LAB, Harvey's Legal Agent Benchmark, agents produce finished legal deliverables graded against dozens of criteria. All-pass counts a task only if it clears them all. Tenet lifts all-pass from 10.8% to 19.7% over the Kimi K3 base, reaching state-of-the-art on LAB Contracts.
Your closed model refused a task it should have done. Ours didn't. DeepSeek V4 Pro on Fireworks: beats Fable 5 on SWE-Bench and LiveCodeBench, a third of the cost per solved task, and it doesn't decline legitimate security-review work. Read more: fireworks.ai/blog/DeepSeekV…
Model reliability is the #1 reason agents fail to complete a task. Frontier models demonstrate capability, then can't deliver consistently at production scale. Sept 2, NYC: Learn to train a real model yourself, then upload it to @Azure Foundry. luma.com/fw.models.nyc
In case you missed it, this is happening tomorrow!
Fireworks@FireworksAI_HQ·Join Fireworks & @LangChain on August 25th for an exclusive hands-on workshop! Learn to build custom eval models for LangSmith and scale your observability stack to supercharge performance. After that? Rooftop happy hour. Let's go: luma.com/phzy304n
In case you missed it, this happening tomorrow!
Fireworks@FireworksAI_HQ·Join Fireworks & @LangChain on August 25th for an exclusive hands-on workshop! Learn to build custom eval models for LangSmith and scale your observability stack to supercharge performance. After that? Rooftop happy hour. Let's go: luma.com/phzy304n
Check out these talks from last month's Nerd Meetup, this time at @posthog's office in SF. @Runloop - Coordinating shared state for multiplayer agents t.co/cYeV5j3Gio - How the architecture of code decides whether AI-written code succeeds Fireworks - Reproducing J-lens on Kimi K3 and Qwen 3.5 Sound like something you'd be interested in? Stay tuned for more info.
We're going global. What are you waiting for? → fireworks.ai/careers
Dimi Nikolaou@diminikola0u·I've joined @FireworksAI_HQ as Field CTO & Europe Lead. Will be building Europe and setting up the FDE Org. Have not been this excited about an industry/space for a long time (ever?) and pumped to get to work. Hiring across every role in London, comment or dm if u want more!
Congrats to the winners of our Kimi K3 Fire Pass Build Contest! - FridgeChef by @seanchenqt - MedRecap by @LotusDecoder - Dayprint by @sethsaler - NotiBridge by @ingwannu1234 - LastLook by @_olegpulatov Check out the entries in the thread below:
Oleg Pulatov@_olegpulatov·I built LastLook for Pi: a second pair of eyes for coding agents. Coding agents have made creating code much faster, but that means verification now takes up a larger share of my time and attention. I built LastLook to get more help from agents with that part of the workflow too. Powered by Kimi K3 Fast on @FireworksAI_HQ, LastLook reviews the exact Git workspace delta and returns a clear PASS/WARN/BLOCK card with concrete suggestions. The review runs outside the primary agent’s context, so it does not add noise to the main conversation. I can send its suggestions directly back into the same Pi chat for the coding agent to address. It is designed to work around the developer’s attention: - Reviews can run automatically while I focus on something else - Results are ready when I return - Reviews can also be requested on demand - Automation and model thinking level are configurable Less manual checking, less context switching, and faster iteration every day. 🔥 #KimiK3onFireworks
Congrats to the winners of our Kimi K3 Fire Pass Build Contest! - FridgeChef by @seanchenqt42 - MedRecap by @LotusDecoder33 - Dayprint by @sethsaler34 - NotiBridge by @ingwannu123435 - LastLook by @_olegpulatov Check out the entries in the thread below:
Oleg Pulatov@_olegpulatov·I built LastLook for Pi: a second pair of eyes for coding agents. Coding agents have made creating code much faster, but that means verification now takes up a larger share of my time and attention. I built LastLook to get more help from agents with that part of the workflow too. Powered by Kimi K3 Fast on @FireworksAI_HQ, LastLook reviews the exact Git workspace delta and returns a clear PASS/WARN/BLOCK card with concrete suggestions. The review runs outside the primary agent’s context, so it does not add noise to the main conversation. I can send its suggestions directly back into the same Pi chat for the coding agent to address. It is designed to work around the developer’s attention: - Reviews can run automatically while I focus on something else - Results are ready when I return - Reviews can also be requested on demand - Automation and model thinking level are configurable Less manual checking, less context switching, and faster iteration every day. 🔥 #KimiK3onFireworks
Sean@seanchenqt·What's for dinner? I let Kimi K3 answer — literally. FridgeChef: snap your fridge, K3 finds the ingredients + plans dinner. The whole app was CODED by K3 too. Demo + try it live: fridgechef.qtkiwi.com @FireworksAI_HQ #KimiK3onFireworks
LotusDecoder@LotusDecoder·#KimiK3onFireworks I built MedRecap, a home medication reconciliation assistant, with Kimi K3 on @FireworksAI_HQ. Upload photos of your pill bottles, prescriptions, and visit notes, and it organizes everything into one medication list, flags discrepancies across the records, and uses FDA drug labels to surface questions worth confirming with a doctor or pharmacist. All demo data is fictional. 👇 I originally planned to build an outdoor hiking guide. Then I discovered a mature product already doing it well, so I switched ideas close to the deadline. The core of MedRecap was built by K3 through Cursor IDE Agent, covering the frontend, backend, image understanding, and report workflow. I spent about $20 on API usage. What impressed me most is how capable open-source models have become at delivering a complete web app, and how stable the @FireworksAI_HQ API was across vision tasks and multi-turn workflows.
Here's a quick look at the live cost monitoring built into Fireworks Nexus, what it tells you, and how to use it. Learn more about Nexus: docs.fireworks.ai/fireworks-nexus
Join us at Glean:GO 2026 at Fort Mason, SF. We'll have two sessions: - Emerging open standards for the AI ecosystem, w/ @aaronamelgar - AI-Native Showcase keynote, w/ Co-Founder & CTO @dzhulgakov Did we mention we'll serve coffee? Register: glean.com/events/glean-g…
Congrats to @harvey on Tenet, their first model and a milestone for legal AI. We’re proud to have co-developed it with them, setting the foundation for firms to build their own specialized intelligence. Read the full writeup: harvey.ai/blog/post-trai…
Trending AND fastest-growing on @tryramp for the 2nd straight month. Four months running on the list. Can confirm @arakharazian's macro take is right: spend is shifting to open-weight, cost-saving inference.
Ara Kharazian@arakharazian·New: Ramp's top SaaS vendors for August 2026 1/ This month, we saw continued growth in model serving / routers as more firms shift spend to cost-saving + open source AI. 2/ A few underrated areas of growth: a lot of competition for AI customer service / sales / voice agents, and seemingly no clear winner in this category given how many different competitors show up on this list month-over-month. 3/ Plus, 9% of firms are using AI to generate images and video for marketing and advertising. That's about half the adoption rate of Figma, and growing quickly.
“After PMF” is the when. How to post-train is more complex: which technique fixes which problem? Why do vibe-based evals break? And how can you cut serving costs once @lqiao broke down the best post-training approach at @sequoia's Own Your Intelligence event.
Sonya Huang 🐥@sonyatweetybird·When should you start post-training your own models? @FireworksAI_HQ CEO @lqiao’s answer: after product-market fit. Not because it's hard... but because only after PMF is the data coming off your product surface worth training on. Lin joined us for our @sequoia "Own Your Intelligence" event to host a workshop on all things post-training; what works, what breaks, and how not to let the model outsmart you. Must listen!! 00:00 Introduction 00:37 What Fireworks sees across thousands of AI applications 02:47 Off-the-shelf APIs and the problem of keeping your taste 03:58 What "owning your intelligence" actually means 05:43 The progression: prompting → RAG → SFT → preferences → RL 07:20 Why this mirrors how humans learn 09:03 Matching the technique to the problem you actually have 10:46 Where teams get stuck: data quality and vibe evals 12:28 Reward hacking: the model that wrote zero lines of code 13:59 Training-to-serving alignment (and why quality drops) 15:55 Post-training in healthcare and security 17:31 From coding to every co-work domain 19:35 Incumbents, cost burden, and not scaling into bankruptcy 21:26 How much control do you want? 23:24 Q&A: What makes a good reward signal 25:00 Q&A: When to start thinking about post-training
Muse Glimmer 30B is now available on Fireworks' Dedicated Training API for both LoRA and Full-Parameter fine-tuning. This is a U.S.-developed, open-weight model and one of the strongest of its size for agentic work, reliable tool use, and long-horizon reasoning. Try it here: t.co/lC3zbN4S7F
Running on AI, powered by caffeine. We're sponsoring Glean GO 2026 in San Francisco. Find us at either of our two coffee stations for an espresso. Then join our Co-Founder & CTO @dzhulgakov on stage for the AI-Native Showcase. See you at Fort Mason! glean.com/events/glean-g….
Thanks to everyone that submitted to the Fire Pass Kimi K3 Community Build Contest! @Prof_OZ our Head of AI Developer Education is testing all of the entries, and highlighting the projects and the builders! First up: FridgeChef by @seanchenqt
Sean@seanchenqt·What's for dinner? I let Kimi K3 answer — literally. FridgeChef: snap your fridge, K3 finds the ingredients + plans dinner. The whole app was CODED by K3 too. Demo + try it live: fridgechef.qtkiwi.com @FireworksAI_HQ #KimiK3onFireworks
In New York early next month? Join Fireworks and @Microsoft September 2nd for an invite only hands-on workshop: train a real model yourself and leave with a customized model that you can upload to Azure Foundry. All you need is your laptop. Easy. RSVP: luma.com/fw.models.nyc
DeepSeek-V4-Pro-0813 is live Day-0 on Fireworks! Built for agentic and tool-augmented workflows, it outperforms Opus 4.8 on quality benchmarks like Terminal Bench 2.1, Cybergym, and DeepSWE at up to 6x the cost efficiency. It is on par with leading frontier models on Vals Index across Coding and Legal benchmarks. What to select: - Use Pro-0813 for heavy reasoning & coding - Use Flash-0731 for budget, less-intensive tasks Start building: t.co/KMngA60KAz
Partnering up with @arcee_ai to help everyone find a route to own their intelligence is easy when their launch videos serve up good vibes. Try our Kimi-K3 on Arcee.
Arcee.ai@arcee_ai·Today we're open-sourcing nac, an agent harness for long-running tasks, and launching the Arcee open models API beta. Nac is available now under Apache 2.0 on GitHub, built for developers running complex, multi-step engineering workloads.
Getting an AI product from prototype to production requires more than choosing a model. This guide from @Microsoft Teams Calls for Startups breaks down how startups can deploy Fireworks models on Microsoft Foundry and build an architecture designed to evolve as they grow. Read the full guide: t.co/uZNWAWi7Mp
Microsoft for Startups@msft4startups·The model you choose for an #MVP can shape cost, latency, flexibility, and future model choice. 🧩 With @FireworksAI_HQ on #MicrosoftFoundry, startups can serve open models in Azure and scale toward production with monitoring, caching, governance, and model experimentation. ➡️ t.co/gA01S039Az #MicrosoftForStartups #GenerativeAI #Developers
We tested out Anthropic’s J-Lens probe at two open models: Kimi K3 and Qwen 3.5-9B. By looking at their internal states, we found that the models are setting up the right vocabulary before a single token hits the page. Check out the research: fireworks.ai/blog/J-Lens-Ki…
A lot of attention is on coding agents, but the more interesting work is in the rest of the stack that actually gets them into production and keeps them reliable. We join @antimetal @braintrust & @browserbase for The AI Dev Stack: Beyond Code. RSVP: luma.com/ai-dev-stack