Skip to content
AI News HubLIVE
Public articles 51Collected articles 54Trust 84Refresh 120 min
Health HealthySource type OfficialFull-text rights Official full textLast ingested 2026-09-25ID fireworks-blogStatus Enabled

Official AI inference and model platform blog; confirm reuse terms before full body display.

Latest public articles

Every byte counts: ARCv3 and the case for cross-region RL

Every byte counts: ARCv3 and the case for cross-region RL Join us for our inaugural conference, Forge 2026 Blog Arcv3 And The Case For Cross Region Rl Every byte counts: ARCv3 and the case for cross-region RL PUBLISHED…

Fireworks AI BlogIn-site articleEvery byte counts: ARCv3 and the case for cross-region RL

Introducing Ember-1

Introducing Ember-1 Join us for our inaugural conference, Forge 2026 Blog Ember 1 Introducing Ember-1 PUBLISHED 9/23/2026 Table of Contents Ember-1: half the tokens, same answers How Fireworks Research built Ember-1 The…

Fireworks AI BlogIn-site articleIntroducing Ember-1

Introducing The Specialized Intelligence Index

Join us for our inaugural conference, Forge 2026 Blog Introducing The Specialized Intelligence Index Introducing The Specialized Intelligence Index PUBLISHED 9/22/2026 Table of Contents The measure of real work Develope…

Fireworks AI BlogIn-site articleIntroducing The Specialized Intelligence Index

The frontier isn’t a model. It’s a router.

The frontier isn’t a model. It’s a router. Join us for our inaugural conference, Forge 2026 Blog The Frontier Isnt A Model Its A Router The frontier isn’t a model. It’s a router. PUBLISHED 9/21/2026 Table of Contents Ho…

Fireworks AI BlogIn-site articleThe frontier isn’t a model. It’s a router.

Phylo brings frontier AI to more scientists with open models on Fireworks

Phylo brings frontier AI to more scientists with open models on Fireworks Join us for our inaugural conference, Forge 2026 Blog Phylo Brings Frontier AI To More Scientists With Open Models On Fireworks Phylo brings fron…

Fireworks AI BlogIn-site articlePhylo brings frontier AI to more scientists with open models on Fireworks

Making the leap to specialized intelligence

Making the leap to specialized intelligence Join us for our inaugural conference, Forge 2026 Blog Making The Leap To Specialized Intelligence Making the leap to specialized intelligence PUBLISHED 9/10/2026 Table of Cont…

Fireworks AI BlogIn-site articleMaking the leap to specialized intelligence

Gen-1 Slides: Opus 5-level decks at a fraction of the cost

Gen-1 Slides: Opus 5-level decks at a fraction of the cost Join us for our inaugural conference, Forge 2026 Blog Gen 1 Slides Opus 5 Level Decks At A Fraction Of The Cost Gen-1 Slides: Opus 5-level decks at a fraction o…

Fireworks AI BlogIn-site articleGen-1 Slides: Opus 5-level decks at a fraction of the cost

Training API now generally available | Fireworks

Training API now generally available | Fireworks Join us for our inaugural conference, Forge 2026 Blog Train Past The Frontier Training API Now Generally Available Train past the frontier: Training API now generally ava…

Fireworks AI BlogIn-site articleTraining API now generally available | Fireworks

DeepSeek V4 Pro: Tops SWE-Bench & Cuts Cost per Task by 3x vs. Fable 5

DeepSeek V4 Pro: Tops SWE-Bench & Cuts Cost per Task by 3x vs. Fable 5 Training API now generally available Blog Deepseekv4pro Fable5 DeepSeek V4 Pro: Tops SWE-Bench & Cuts Cost per Task by 3x vs. Fable 5 PUBLISHED 8/26…

Fireworks AI BlogIn-site articleDeepSeek V4 Pro: Tops SWE-Bench & Cuts Cost per Task by 3x vs. Fable 5

Post-training Kimi K3 with Harvey for long-horizon legal work

Post-training Kimi K3 with Harvey for long-horizon legal work DeepSeek-V4-Pro-0813 available now on Fireworks Blog Post Training Kimi K3 With Harvey For Long Horizon Legal Work Post-training Kimi K3 with Harvey for long…

Fireworks AI BlogIn-site articlePost-training Kimi K3 with Harvey for long-horizon legal work

Fireworks AI

Fireworks AI DeepSeek-V4-Pro-0813 available now on Fireworks Blog Deepseek V4 Pro Security DeepSeek V4 Pro is Redefining Security Agent Economics PUBLISHED 8/26/2026 TLDR; DeepSeek V4 Pro 0813 recorded zero refusals acr…

Fireworks AI BlogIn-site articleFireworks AI

Can open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on Kimi K3 & Qwen3.5-9B

DeepSeek-V4-Pro-0813 available now on Fireworks Blog J Lens Kimi K3 Qwen Can open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on Kimi K3 & Qwen3.5-9B PUBLISHED 8/12/2026 Table of…

Fireworks AI BlogIn-site articleCan open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on Kimi K3 & Qwen3.5-9B

Fireworks AI

Fireworks AI Kimi K3 on Fireworks: Frontier Intelligence You Can Own Blog Meta Muse Glimmer Muse Glimmer from Meta on Fireworks: Ideal for your Always-On Agents PUBLISHED 8/10/2026 Table of Contents Inside the Architect…

Fireworks AI BlogIn-site articleFireworks AI

Your AI Performance Stack is Fireworks Models with Voyage AI embeddings

Your AI Performance Stack is Fireworks Models with Voyage AI embeddings Kimi K3 on Fireworks: Frontier Intelligence You Can Own Blog Voyage AI Models Now On Fireworks Your AI performance stack is Fireworks + Voyage AI P…

Fireworks AI BlogIn-site articleYour AI Performance Stack is Fireworks Models with Voyage AI embeddings

Three Tests to Run Before You Switch from LoRA to FullFT

Three Tests to Run Before You Switch from LoRA to FullFT Kimi K3 on Fireworks: Frontier Intelligence You Can Own Blog Three Tests To Run Before You Switch From Lora To Fullft Three Tests to Run Before You Switch from Lo…

Fireworks AI BlogIn-site articleThree Tests to Run Before You Switch from LoRA to FullFT

Trilogy’s Playbook for Open-Weight Cybersecurity with Kimi K3

Trilogy's AI Center of Excellence has released a cybersecurity playbook built around Kimi K3 on Fireworks. The playbook emphasizes the advantages of open-weight models for defensive AI, enabling continuous high-volume security workflows without dependency on gated, expensive models. It details a practical reference stack using deterministic tools for structure and Kimi K3 for judgment, with a path-centric audit approach. Fireworks provides managed inference, allowing teams to scale independently while retaining control over their audit workflows.

Fireworks AI BlogIn-site articleTrilogy’s Playbook for Open-Weight Cybersecurity with Kimi K3

Fireworks Nexus: Drop-in Open Frontier Intelligence for Teams with Budgets

Fireworks AI launches Fireworks Nexus to help engineering teams reduce AI costs by routing routine tasks to cost-effective open-source models without changing workflows. Early tests show a third reduction in cost per merged PR and blended token rates about a quarter of closed model labs.

Fireworks AI BlogIn-site articleFireworks Nexus: Drop-in Open Frontier Intelligence for Teams with Budgets

Fireworks AI Launches LoRA Training for Kimi K3: Frontier Intelligence at Your Fingertips

Fireworks AI introduces serverless LoRA training for its massive 2.8 trillion parameter MoE model, Kimi K3. LoRA adapters enable cost-effective fine-tuning with pay-per-token pricing and flexible serving. Two example tasks demonstrate how small adapters can teach K3 new objectives or tool-use loops within minutes. The article emphasizes the critical role of reward design and the data flywheel effect.

Fireworks AI BlogIn-site articleFireworks AI Launches LoRA Training for Kimi K3: Frontier Intelligence at Your Fingertips

Kimi K3 on Fireworks: Frontier Intelligence You Can Own

Kimi K3, an open-weight frontier model with 2.8 trillion parameters, is now available on Fireworks for inference and training. It rivals top closed models like Fable 5, Opus 5, and GPT 5.5, excelling in coding, writing, legal analysis, and multimodal tasks. Fireworks offers US-hosted inference with zero data retention and pay-per-token pricing, with options for serverless and batch inference, as well as serverless training.

Fireworks AI BlogIn-site articleKimi K3 on Fireworks: Frontier Intelligence You Can Own

Best Open Source LLMs in 2026: We Reviewed 7 Models

The article reviews nine open-source LLMs as of mid-2026, comparing them on benchmarks, context windows, modality support, and licensing. It highlights Kimi K3 as the top performer but notes its weights are pending, and recommends GLM 5.2 as the best currently available open model on Fireworks. Other models like DeepSeek-V4-Pro, MiniMax M3, and gpt-oss-120b are evaluated for specific use cases.

Fireworks AI BlogIn-site articleBest Open Source LLMs in 2026: We Reviewed 7 Models

Heidi x Fireworks: Bridging the Gap in Frontier Model Performance

Heidi Health partners with Fireworks AI to surpass proprietary frontier model quality using SFT and RFT, achieving 3.5x lower latency, 7-second inference time, and significant cost savings. Success hinges on aggressive data filtering and large batch sizes (1.5M effective tokens), enabling open models to rival frontier models.

Fireworks AI BlogIn-site articleHeidi x Fireworks: Bridging the Gap in Frontier Model Performance

Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.

Fireworks AI announces research showing open-source Kimi K3 is competitive with closed-source Fable 5 across ~1,000 agentic tasks. Routing between them achieves 93% accuracy and up to 50x cost savings, suggesting the best AI comes from a mixture of models.

Fireworks AI BlogIn-site articleKimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.

Optimizing MiniMax M3 Sparse Attention on NVIDIA Blackwell

Fireworks AI built a Blackwell (SM100) kernel for MiniMax M3 sparse attention that uses a KV-stationary execution path, loading each KV block once, achieving ~980 TFLOP/s throughput and ~4.1 TB/s HBM bandwidth, with 1.9–2.4× speedup over a query-stationary baseline (FlashInfer) and ~1.6× over open-source MSA kernels. The article details the algorithm, kernel design, optimizations, and I/O cost model.

Fireworks AI BlogIn-site articleOptimizing MiniMax M3 Sparse Attention on NVIDIA Blackwell

Fireworks Secures $1.5 Billion in Series D Funding

Fireworks announces a $1.505 billion Series D at a $17.5 billion valuation, surpassing $1 billion in annualized revenue and processing over 40 trillion tokens daily, with 95% from specialized customer models.

Fireworks AI BlogIn-site articleFireworks Secures $1.5 Billion in Series D Funding

How Gumloop Scaled Open-Weight Model Usage 7x in 3 Weeks with Fireworks AI

Gumloop achieved 7x growth in open-weight model agent chats within three weeks by partnering with Fireworks AI, saving up to 72% on costs while maintaining quality.

Fireworks AI BlogIn-site articleHow Gumloop Scaled Open-Weight Model Usage 7x in 3 Weeks with Fireworks AI

Open, frontier, and yours: LangChain Deep Agents on NVIDIA Nemotron 3 Ultra, running on Fireworks

LangChain has tuned its Deep Agents harness for NVIDIA Nemotron 3 Ultra, achieving benchmark-leading agent performance among open models at 10x lower cost than closed alternatives. The tuned harness is available in LangChain Deep Agents, and Nemotron 3 Ultra runs on Fireworks with day-zero support.

Fireworks AI BlogIn-site articleOpen, frontier, and yours: LangChain Deep Agents on NVIDIA Nemotron 3 Ultra, running on Fireworks

How I shipped a month of engineering work in four days with GLM 5.2 Fast

Using GLM 5.2 Fast via FireConnect on Claude Code, the author designed, planned, and implemented a GPU scheduler reclaim feature in four days at a cost of $218 in inference tokens—a task normally scoped at one engineer-month. The article highlights how fast inference speed (400 tokens/sec) eliminated context switching, low token cost removed usage anxiety, and model quality handled complex concurrent logic.

Fireworks AI BlogIn-site articleHow I shipped a month of engineering work in four days with GLM 5.2 Fast

GLM 5.2 Fast is live on Fireworks

GLM 5.2 Fast is available on Fireworks, offering Opus-level intelligence at open-source rates. It runs 2-3x faster than Standard on shared serverless infrastructure, with a full 1M-token context window, aggressive prompt caching, and structured output support. Fireworks achieved 446 tok/sec and optimized the model's MoE and sparse attention architectures for real-world workloads.

Fireworks AI BlogIn-site articleGLM 5.2 Fast is live on Fireworks

How Factory Grew Open Model Usage 2-3x in Six Months on Fireworks

Factory is building agent-native software development. Its agents, Droids, cover the entire SDLC. By using Fireworks for open-weight model inference, Factory increased open model usage 2-3x in six months, cutting task costs to 6-20% of frontier models and boosting throughput 5-15x per dollar. This is due to Fireworks' complete model coverage, day-zero availability, high reliability, and low latency.

Fireworks AI BlogIn-site articleHow Factory Grew Open Model Usage 2-3x in Six Months on Fireworks

Cursor Composer 2 + Fireworks AI

Cursor released Composer 2, a coding model optimized for the Cursor development environment. Based on Kimi 2.5, it combines continual pre-training and large-scale reinforcement learning to achieve frontier-level coding performance while reducing inference cost by 6-10x. Fireworks AI provides the distributed inference infrastructure to make RL scalable.

Fireworks AI BlogIn-site articleCursor Composer 2 + Fireworks AI

All sources