Every byte counts: ARCv3 and the case for cross-region RL Join us for our inaugural conference, Forge 2026 Blog Arcv3 And The Case For Cross Region Rl Every byte counts: ARCv3 and the case for cross-region RL PUBLISHED…
Introducing Ember-1 Join us for our inaugural conference, Forge 2026 Blog Ember 1 Introducing Ember-1 PUBLISHED 9/23/2026 Table of Contents Ember-1: half the tokens, same answers How Fireworks Research built Ember-1 The…
Join us for our inaugural conference, Forge 2026 Blog Introducing The Specialized Intelligence Index Introducing The Specialized Intelligence Index PUBLISHED 9/22/2026 Table of Contents The measure of real work Develope…
The frontier isn’t a model. It’s a router. Join us for our inaugural conference, Forge 2026 Blog The Frontier Isnt A Model Its A Router The frontier isn’t a model. It’s a router. PUBLISHED 9/21/2026 Table of Contents Ho…
Phylo brings frontier AI to more scientists with open models on Fireworks Join us for our inaugural conference, Forge 2026 Blog Phylo Brings Frontier AI To More Scientists With Open Models On Fireworks Phylo brings fron…
Making the leap to specialized intelligence Join us for our inaugural conference, Forge 2026 Blog Making The Leap To Specialized Intelligence Making the leap to specialized intelligence PUBLISHED 9/10/2026 Table of Cont…
Gen-1 Slides: Opus 5-level decks at a fraction of the cost Join us for our inaugural conference, Forge 2026 Blog Gen 1 Slides Opus 5 Level Decks At A Fraction Of The Cost Gen-1 Slides: Opus 5-level decks at a fraction o…
Training API now generally available | Fireworks Join us for our inaugural conference, Forge 2026 Blog Train Past The Frontier Training API Now Generally Available Train past the frontier: Training API now generally ava…
DeepSeek V4 Pro: Tops SWE-Bench & Cuts Cost per Task by 3x vs. Fable 5 Training API now generally available Blog Deepseekv4pro Fable5 DeepSeek V4 Pro: Tops SWE-Bench & Cuts Cost per Task by 3x vs. Fable 5 PUBLISHED 8/26…
Post-training Kimi K3 with Harvey for long-horizon legal work DeepSeek-V4-Pro-0813 available now on Fireworks Blog Post Training Kimi K3 With Harvey For Long Horizon Legal Work Post-training Kimi K3 with Harvey for long…
Fireworks AI DeepSeek-V4-Pro-0813 available now on Fireworks Blog Deepseek V4 Pro Security DeepSeek V4 Pro is Redefining Security Agent Economics PUBLISHED 8/26/2026 TLDR; DeepSeek V4 Pro 0813 recorded zero refusals acr…
DeepSeek-V4-Pro-0813 available now on Fireworks Blog J Lens Kimi K3 Qwen Can open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on Kimi K3 & Qwen3.5-9B PUBLISHED 8/12/2026 Table of…
Fireworks AI Kimi K3 on Fireworks: Frontier Intelligence You Can Own Blog Meta Muse Glimmer Muse Glimmer from Meta on Fireworks: Ideal for your Always-On Agents PUBLISHED 8/10/2026 Table of Contents Inside the Architect…
Your AI Performance Stack is Fireworks Models with Voyage AI embeddings Kimi K3 on Fireworks: Frontier Intelligence You Can Own Blog Voyage AI Models Now On Fireworks Your AI performance stack is Fireworks + Voyage AI P…
Three Tests to Run Before You Switch from LoRA to FullFT Kimi K3 on Fireworks: Frontier Intelligence You Can Own Blog Three Tests To Run Before You Switch From Lora To Fullft Three Tests to Run Before You Switch from Lo…
Trilogy's AI Center of Excellence has released a cybersecurity playbook built around Kimi K3 on Fireworks. The playbook emphasizes the advantages of open-weight models for defensive AI, enabling continuous high-volume security workflows without dependency on gated, expensive models. It details a practical reference stack using deterministic tools for structure and Kimi K3 for judgment, with a path-centric audit approach. Fireworks provides managed inference, allowing teams to scale independently while retaining control over their audit workflows.
Fireworks AI launches Fireworks Nexus to help engineering teams reduce AI costs by routing routine tasks to cost-effective open-source models without changing workflows. Early tests show a third reduction in cost per merged PR and blended token rates about a quarter of closed model labs.
Fireworks AI introduces serverless LoRA training for its massive 2.8 trillion parameter MoE model, Kimi K3. LoRA adapters enable cost-effective fine-tuning with pay-per-token pricing and flexible serving. Two example tasks demonstrate how small adapters can teach K3 new objectives or tool-use loops within minutes. The article emphasizes the critical role of reward design and the data flywheel effect.
Kimi K3, an open-weight frontier model with 2.8 trillion parameters, is now available on Fireworks for inference and training. It rivals top closed models like Fable 5, Opus 5, and GPT 5.5, excelling in coding, writing, legal analysis, and multimodal tasks. Fireworks offers US-hosted inference with zero data retention and pay-per-token pricing, with options for serverless and batch inference, as well as serverless training.
The article reviews nine open-source LLMs as of mid-2026, comparing them on benchmarks, context windows, modality support, and licensing. It highlights Kimi K3 as the top performer but notes its weights are pending, and recommends GLM 5.2 as the best currently available open model on Fireworks. Other models like DeepSeek-V4-Pro, MiniMax M3, and gpt-oss-120b are evaluated for specific use cases.
Heidi Health partners with Fireworks AI to surpass proprietary frontier model quality using SFT and RFT, achieving 3.5x lower latency, 7-second inference time, and significant cost savings. Success hinges on aggressive data filtering and large batch sizes (1.5M effective tokens), enabling open models to rival frontier models.
Fireworks AI announces research showing open-source Kimi K3 is competitive with closed-source Fable 5 across ~1,000 agentic tasks. Routing between them achieves 93% accuracy and up to 50x cost savings, suggesting the best AI comes from a mixture of models.
Fireworks AI built a Blackwell (SM100) kernel for MiniMax M3 sparse attention that uses a KV-stationary execution path, loading each KV block once, achieving ~980 TFLOP/s throughput and ~4.1 TB/s HBM bandwidth, with 1.9–2.4× speedup over a query-stationary baseline (FlashInfer) and ~1.6× over open-source MSA kernels. The article details the algorithm, kernel design, optimizations, and I/O cost model.
Fireworks announces a $1.505 billion Series D at a $17.5 billion valuation, surpassing $1 billion in annualized revenue and processing over 40 trillion tokens daily, with 95% from specialized customer models.
Gumloop achieved 7x growth in open-weight model agent chats within three weeks by partnering with Fireworks AI, saving up to 72% on costs while maintaining quality.
LangChain has tuned its Deep Agents harness for NVIDIA Nemotron 3 Ultra, achieving benchmark-leading agent performance among open models at 10x lower cost than closed alternatives. The tuned harness is available in LangChain Deep Agents, and Nemotron 3 Ultra runs on Fireworks with day-zero support.
Using GLM 5.2 Fast via FireConnect on Claude Code, the author designed, planned, and implemented a GPU scheduler reclaim feature in four days at a cost of $218 in inference tokens—a task normally scoped at one engineer-month. The article highlights how fast inference speed (400 tokens/sec) eliminated context switching, low token cost removed usage anxiety, and model quality handled complex concurrent logic.
GLM 5.2 Fast is available on Fireworks, offering Opus-level intelligence at open-source rates. It runs 2-3x faster than Standard on shared serverless infrastructure, with a full 1M-token context window, aggressive prompt caching, and structured output support. Fireworks achieved 446 tok/sec and optimized the model's MoE and sparse attention architectures for real-world workloads.
Factory is building agent-native software development. Its agents, Droids, cover the entire SDLC. By using Fireworks for open-weight model inference, Factory increased open model usage 2-3x in six months, cutting task costs to 6-20% of frontier models and boosting throughput 5-15x per dollar. This is due to Fireworks' complete model coverage, day-zero availability, high reliability, and low latency.
Cursor released Composer 2, a coding model optimized for the Cursor development environment. Based on Kimi 2.5, it combines continual pre-training and large-scale reinforcement learning to achieve frontier-level coding performance while reducing inference cost by 6-10x. Fireworks AI provides the distributed inference infrastructure to make RL scalable.