Three Tests to Run Before You Switch from LoRA to FullFT Kimi K3 on Fireworks: Frontier Intelligence You Can Own Blog Three Tests To Run Before You Switch From Lora To Fullft Three Tests to Run Before You Switch from Lo…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Three Tests to Run Before You Switch from LoRA to FullFT Kimi K3 on Fireworks: Frontier Intelligence You Can Own Blog Three Tests To Run Before You Switch From Lora To Fullft Thre…
Trilogy's AI Center of Excellence has released a cybersecurity playbook built around Kimi K3 on Fireworks. The playbook emphasizes the advantages of open-weight models for defensive AI, enabling continuous high-volume security workflows without dependency on gated, expensive models. It details a practical reference stack using deterministic tools for structure and Kimi K3 for judgment, with a path-centric audit approach. Fireworks provides managed inference, allowing teams to scale independently while retaining control over their audit workflows.
Trilogy's playbook uses Kimi K3 on Fireworks for open-weight cybersecurity, emphasizing deployment flexibility.
The playbook separates deterministic tools (indexing, parsing) from AI judgment (Kimi K3) for efficient security audits.
Fireworks AI launches Fireworks Nexus to help engineering teams reduce AI costs by routing routine tasks to cost-effective open-source models without changing workflows. Early tests show a third reduction in cost per merged PR and blended token rates about a quarter of closed model labs.
Fireworks Nexus provides enterprise controls, cost observability, and workflow continuity, enabling centralized AI management.
FireConnect installs with one line, seamlessly connecting existing tools like Claude Code and Codex to open models.
Fireworks AI introduces serverless LoRA training for its massive 2.8 trillion parameter MoE model, Kimi K3. LoRA adapters enable cost-effective fine-tuning with pay-per-token pricing and flexible serving. Two example tasks demonstrate how small adapters can teach K3 new objectives or tool-use loops within minutes. The article emphasizes the critical role of reward design and the data flywheel effect.
Serverless LoRA training for Kimi K3 is now available, charging per token with no cluster management needed.
LoRA adapters are tiny (MBs), allowing cheap fine-tuning for specific behaviors, and can be merged live or served in multi-LoRA setups.
Kimi K3, an open-weight frontier model with 2.8 trillion parameters, is now available on Fireworks for inference and training. It rivals top closed models like Fable 5, Opus 5, and GPT 5.5, excelling in coding, writing, legal analysis, and multimodal tasks. Fireworks offers US-hosted inference with zero data retention and pay-per-token pricing, with options for serverless and batch inference, as well as serverless training.
Kimi K3 is the first open model to reach 2.8T parameters, achieving #1 in front-end code and top rankings in various benchmarks.
It features Kimi Delta Attention and MoE architecture for cost-efficient performance with a low activation rate.
The article reviews nine open-source LLMs as of mid-2026, comparing them on benchmarks, context windows, modality support, and licensing. It highlights Kimi K3 as the top performer but notes its weights are pending, and recommends GLM 5.2 as the best currently available open model on Fireworks. Other models like DeepSeek-V4-Pro, MiniMax M3, and gpt-oss-120b are evaluated for specific use cases.
Kimi K3 leads benchmarks but full weights are pending release in July 2026.
GLM 5.2 is the best open model available on Fireworks with a 1,040k context window.
Heidi Health partners with Fireworks AI to surpass proprietary frontier model quality using SFT and RFT, achieving 3.5x lower latency, 7-second inference time, and significant cost savings. Success hinges on aggressive data filtering and large batch sizes (1.5M effective tokens), enabling open models to rival frontier models.
Heidi partners with Fireworks to shift from closed to open models for better performance and cost savings.
Using SFT and RFT, model quality surpasses frontier models like Gemini Pro in internal evaluations.
Fireworks AI announces research showing open-source Kimi K3 is competitive with closed-source Fable 5 across ~1,000 agentic tasks. Routing between them achieves 93% accuracy and up to 50x cost savings, suggesting the best AI comes from a mixture of models.
Kimi K3 and Fable 5 perform similarly overall but have different strengths across task domains.
Oracle routing selects K3 for 72-96% of tasks, outperforming any single model.
Fireworks AI built a Blackwell (SM100) kernel for MiniMax M3 sparse attention that uses a KV-stationary execution path, loading each KV block once, achieving ~980 TFLOP/s throughput and ~4.1 TB/s HBM bandwidth, with 1.9–2.4× speedup over a query-stationary baseline (FlashInfer) and ~1.6× over open-source MSA kernels. The article details the algorithm, kernel design, optimizations, and I/O cost model.
Sparse attention reduces compute and memory costs for long-context inference, but data-dependent selection leads to irregular memory access, negating theoretical speedups.
Fireworks AI's KV-stationary kernel loads each selected KV block once in the outer loop and attends all queries that chose it in the inner loop, enabling efficient compute.
Fireworks announces a $1.505 billion Series D at a $17.5 billion valuation, surpassing $1 billion in annualized revenue and processing over 40 trillion tokens daily, with 95% from specialized customer models.
Fireworks raises $1.505B in Series D, valuation at $17.5B.
Company exceeds $1B ARR, processes 40T+ tokens daily.
Gumloop achieved 7x growth in open-weight model agent chats within three weeks by partnering with Fireworks AI, saving up to 72% on costs while maintaining quality.
Gumloop migrated internal agent from Opus 4.8 to GLM-5.2 with no user-noticeable difference.
Achieved 72% cost savings and 7x growth in open-weight model usage.
LangChain has tuned its Deep Agents harness for NVIDIA Nemotron 3 Ultra, achieving benchmark-leading agent performance among open models at 10x lower cost than closed alternatives. The tuned harness is available in LangChain Deep Agents, and Nemotron 3 Ultra runs on Fireworks with day-zero support.
LangChain Deep Agents on Nemotron 3 Ultra deliver frontier performance at 10x lower cost per task.
Fireworks enables post-training customization, allowing enterprises to own and improve their models.
Using GLM 5.2 Fast via FireConnect on Claude Code, the author designed, planned, and implemented a GPU scheduler reclaim feature in four days at a cost of $218 in inference tokens—a task normally scoped at one engineer-month. The article highlights how fast inference speed (400 tokens/sec) eliminated context switching, low token cost removed usage anxiety, and model quality handled complex concurrent logic.
Used GLM 5.2 Fast via FireConnect on Claude Code to build a GPU scheduler reclaim capability. Delivered 4 PRs, ~3000 lines of code, 34 passing tests in 4 days at $218 inference cost.
Fast inference (~400 tokens/sec) enabled real-time design collaboration without context switching, compressing the design phase from weeks to one day.
GLM 5.2 Fast is available on Fireworks, offering Opus-level intelligence at open-source rates. It runs 2-3x faster than Standard on shared serverless infrastructure, with a full 1M-token context window, aggressive prompt caching, and structured output support. Fireworks achieved 446 tok/sec and optimized the model's MoE and sparse attention architectures for real-world workloads.
GLM 5.2 Fast is 2-3x faster than Standard, with 1M-token context and prompt caching at $0.14 per 1M cached tokens.
Fireworks optimized MoE and sparse attention separately for better parallelism and memory efficiency.
Factory is building agent-native software development. Its agents, Droids, cover the entire SDLC. By using Fireworks for open-weight model inference, Factory increased open model usage 2-3x in six months, cutting task costs to 6-20% of frontier models and boosting throughput 5-15x per dollar. This is due to Fireworks' complete model coverage, day-zero availability, high reliability, and low latency.
Factory achieved day-zero access to all open models via Fireworks, driving 2-3x usage growth.
Open models cost 6-20% of frontier models while maintaining quality, making automation economical.
Cursor released Composer 2, a coding model optimized for the Cursor development environment. Based on Kimi 2.5, it combines continual pre-training and large-scale reinforcement learning to achieve frontier-level coding performance while reducing inference cost by 6-10x. Fireworks AI provides the distributed inference infrastructure to make RL scalable.
Composer 2 is a specialized coding model for Cursor's environment, improved through continual pre-training and reinforcement learning.
It achieves top scores on CursorBench, Terminal-Bench, and SWE-bench Multilingual.
This article presents a worker-advisor architecture that combines open-source worker agents with a closed-source advisor model, achieving near-frontier performance on multiple benchmarks at significantly lower costs. The GLM-5.2 + Opus 4.8 combination shows consistent improvements across SWE-bench Pro, Terminal-Bench 2.1, and Legal Agent Bench, with cost savings of 19% to 67% compared to using Opus alone as the worker.
An open-source worker (Kimi-K2.6 or GLM-5.2) drives the task end-to-end, consulting a closed-source frontier model (Claude Opus 4.8) once for review.
Lifts of +4 to +7 pp on SWE-bench Pro, +4 to +8 pp on Terminal-Bench 2.1, and +1 to +4 pp on Legal Agent Bench.
Fireworks AI is migrating all self-serve accounts to prepaid billing effective July 1, 2026. Users can switch now to control the timing or be automatically migrated. Prepaid billing offers predictability with credit purchases and auto-reload options. Contracted customers are not affected.
Fireworks AI moves to prepaid billing for self-serve accounts on July 1, 2026.
Users can either switch now or be migrated automatically on that date.
GLM 5.2, the latest open-source model from Z.ai (formerly Zhipu), is now available on Fireworks inference platform. It leads coding benchmarks, features a 1M-token context window for long-horizon tasks, and is MIT-licensed. Fireworks validates performance independently, emphasizing infrastructure over routing.
GLM 5.2 is now live on Fireworks inference, day zero.
It is the strongest open-source model for coding, with a 1M-token context.
Moonshot AI has released Kimi K2.7 Code, the latest coding model in the K2 line, now available on Fireworks AI with Day-0 support. The model uses 30% fewer reasoning tokens than its predecessor while achieving higher scores on coding benchmarks. This reduction in reasoning tokens significantly lowers the cost per completed task for agentic workflows. Fireworks offers three serving tiers: Standard, Priority, and Fast (coming soon), catering to different reliability and speed needs.
Kimi K2.7 Code uses 30% fewer reasoning tokens than K2.6 but scores higher on coding evals.
Lower reasoning tokens reduce overall cost per task in agentic workflows due to compounding effects.
Alibaba has partnered with Fireworks to host Qwen 3.7 Plus on its infrastructure, making the flagship multimodal model available via serverless API. Designed for agent loops, it supports thinking and non-thinking modes, a 262K context window, and offers a 50% price reduction over its predecessor. Fireworks provides direct inference with zero data retention and 99.9% uptime SLA.
Qwen 3.7 Plus is now available exclusively on Fireworks as a serverless API.
The model is built for agentic workflows, supporting image input and reasoning preservation across turns.
MiniMax releases flagship model M3 with over 500K token context window, native multimodality, and MiniMax Sparse Attention (MSA) architecture, delivering frontier-level coding and agentic capabilities at a fraction of the cost of previous models.
MiniMax M3 supports over 500K token context, expanding to 1M soon.
Uses MiniMax Sparse Attention (MSA) for sub-quadratic scaling, 4x faster than alternatives.
NVIDIA's Nemotron 3 Ultra, an open model optimized for long-running autonomous agents, launches with day-zero support on Fireworks. With 550B total parameters, hybrid Transformer-Mamba MoE architecture, and up to 1M context, it offers 5x faster inference and 30% lower cost for agentic tasks compared to other open models. Fireworks provides dedicated GPU deployments and a unified platform for training and inference.
Nemotron 3 Ultra is an open model designed for long-running autonomous agents, featuring 550B total and 55B active parameters.
It uses a hybrid Transformer-Mamba MoE architecture with up to 1M context length.
Fireworks AI and Harvey explore two system-level techniques on Legal Agent Benchmark (LAB) to reduce reliance on single frontier model calls while achieving frontier-level performance at lower cost. A hybrid harness with open-source GLM 5.1 worker and Claude Opus 4.7 advisor achieves 18/100 all-pass at $368, surpassing Opus alone (14/100 at $954). Post-training of Kimi K2.6 via SFT and RFT yields 15/100 all-pass at $84 and improved mean scores respectively.
Hybrid harness with open-source worker and frontier advisor as callable tool achieves higher all-pass at lower cost than end-to-end frontier model.
Post-training on Fireworks: SFT lifts all-pass from 11 to 15/100; RFT boosts mean score from 0.863 to 0.886.
Trilogy's AI Center of Excellence evaluated Fireworks AI as inference infrastructure to standardize open-weight model usage, reducing costs and enabling billion-token-scale agentic workflows.
Trilogy adopted Fireworks AI as the inference layer for enterprise open-weight models.
Reduced cost to ~1/5 of proprietary systems and eliminated rate limit issues.
A benchmark of 720 browser agent tasks reveals that structured output reliability, not raw intelligence, is the bottleneck in agentic AI. Gemini 2.5 Flash incurred a 22.9% execution tax due to malformed JSON, while Kimi K2.5 had zero. This tax compounds into higher latency, cost, and failure rates. The report introduces Reliability-Adjusted Accuracy and cost-per-successful-task metrics.
Agent Execution Tax measures wasted inference from structured output failures; top model had 22.9% tax.
Gemini 2.5 Flash had 86.7% probability of at least one parse retry per task; Kimi K2.5 had 0%.
Fireworks AI launches Serverless 2.0, offering Standard, Priority, and Fast inference paths through a single API without reserved capacity. The Priority path provides stronger request admission under congestion, while the Fast path delivers roughly 2x throughput. The update also clarifies error codes by separating load shedding (503) from rate limits (429), improving retry logic and alerting.
Serverless 2.0 introduces three serving intents: Standard (default), Priority (stronger admission under load), and Fast (higher token throughput).
Priority achieved 0% 503 error rate in peak-load testing versus 0.082% for Standard.
As a Tier 1 AWS Premier Partner, Innovative Solutions transformed its services delivery by migrating its inference layer to Fireworks AI. The DarcyIQ platform evolved from an internal productivity tool into a multi-agent execution system, compressing contract cycles from 30–45 days to ~3 days, doubling delivery throughput, and making inference costs predictable and controllable.
Innovative Solutions migrated its inference layer from Anthropic to Fireworks AI, reducing model integration overhead and achieving stable, cost-predictable inference.
DarcyIQ evolved into a multi-agent execution system covering sales, scoping, and delivery, cutting contract cycles to ~3 days.
Fireworks AI has acquired Hathora, a company specializing in low-latency container orchestration for gaming, to enhance its AI inference platform with millisecond-level routing and global scalability.
Fireworks AI acquires Hathora to improve AI inference latency.
Hathora's orchestration handles routing across 14 regions and multiple clouds.
Fireworks AI announces public preview of its high-performance open model inference on Microsoft Foundry, integrating leading models like DeepSeek V3.2 and Kimi K2.5 into Azure with BYOW and flexible pricing.
Fireworks AI on Microsoft Foundry brings best-in-class open model inference to Azure. Available models include DeepSeek V3.2, Kimi K2.5, and more. Supports bring-your-own-weights and serverless or provisioned throughput pricing.