Fireworks Nexus: Drop-in Open Frontier Intelligence for Teams with Budgets
Fireworks AI launches Fireworks Nexus to help engineering teams reduce AI costs by routing routine tasks to cost-effective open-source models without changing workflows. Early tests show a third reduction in cost per merged PR and blended token rates about a quarter of closed model labs.
Fireworks AI
Kimi K3 on Fireworks: Frontier Intelligence You Can Own
Blog
Fireworks Nexus
Fireworks Nexus: Drop-in Open Frontier Intelligence for Teams with Budgets
PUBLISHED 7/26/2026
Open models just crossed the intelligence-cost curve. Most of your team's bill is from routine work at frontier prices. We make it easy to start optimizing your AI spend, without changing how your engineers work.
If you run engineering at any real scale, you've already had this conversation: someone in finance asks what the AI cost is going to be next quarter, and the honest answer is that nobody knows.
When Forbes reported that Uber had spent its entire 2026 AI budget by April, the news made global headlines. They rolled out Claude Code to roughly 5,000 engineers in December and watched agentic adoption climb from about a third of engineers to more than four-fifths inside two months. The CTO reported burning $1,200 in a single two-hour session. Uber’s COO said publicly that the link between that spend and shipped customer value is hard to draw.
But something’s been quietly changing: open models are now reliable for the vast majority of real tasks. We found most organizations were running routine work at frontier prices, but due to the operational complexity of running open-weight models at scale, it’s been hard for them to switch. So we took all the background work out of running the best, curated open models like Kimi-K3 and GLM-5.2 at scale, and made it something you drop straight into your existing workflow. We’ve been already testing Fireworks Nexus with teams like Notion and Doximity where the preliminary results show that Fireworks Nexus cuts a third off per merged PR costs, and has a blended token rate roughly a quarter of the closed model labs.
Available Today for Engineering Organizations: Fireworks Nexus
Fireworks Nexus connects the AI tools your developers already use to a managed layer of open-weight models, with enterprise controls and intelligent routing. Engineering teams gain visibility and control over the AI powering their workflows while lowering costs and maintaining the developer experience they already know.
It is composed of three main components:
- Enterprise Controls & Cost Observability
Fireworks Nexus gives engineering and IT teams centralized control over how AI is used across the organization. Set budgets at the team or company level, track ROI across models and tools, and enforce policies from a single place. Behind the scenes, every request runs on the same production inference platform trusted by leading AI-native companies, with US-hosted endpoints, zero data retention, and enterprise-grade performance across 20 global data centers.
Figure 1: Screenshot of a Sample Fireworks Nexus Portal
- Workflow Continuity
FireConnect is a one-line install that maps appropriate models based on harness configurations intuitively. Keep using your current tools and harnesses like Claude Code, Codex, OpenCode. It ensures developers face no disruption, while benefiting from optimized infrastructure including higher cache rates that can drive material cost savings. FireConnect works through our Fireworks Serverless APIs that are Anthropic and OpenAI API compatible, so most tools connect with a base URL and a model ID. To get started, FireConnect is open sourced under Apache 2.0 and can be installed from the Fireworks Dashboard in a single command.
Figure 2: FireConnect API Key
- Intelligent Traffic Management and Migration
We built a custom routing endpoint for organizations with existing AI contracts. A custom trained model scores each request's difficulty. Requests that are routine tasks go to a cost-effective open-weight model served by Fireworks, difficult tasks pass through to your existing provider on your own key, which is never stored server-side. It serves as an intermediary layer that dynamically routes tasks based on complexity and optimal caching. This typically delivers a 3–5x cost reduction without sacrificing quality. This feature is in research preview and today routes between Claude Opus 5 and GLM 5.2, so the pass-through path needs an Anthropic key. Want to stay entirely on open weight models? We also route between K3 and GLM 5.2
Figure 3: FireRouter Component OverviewFigure 4: FireRouter Routing Preference
Proven on Real Engineering Work
The value of Fireworks Nexus isn’t just lower token prices. It’s getting the same or better engineering outcomes while spending less.
Independent evaluations back this up. Faros ran 211 real engineering tasks from customer repositories across seven model-and-harness combinations and published the full methodology. Claude Code running on GLM-5.2 slightly outperformed Claude Code on Opus 4.8 while costing roughly half as much per completed task. Arize evaluated 2,400 agent runs and found that frontier-priced models offered little advantage on routine work, while intelligent routing preserved their strengths where they mattered most. Their evaluation harness is open source, so you can run the same analysis on your own workloads.
Fireworks Nexus gives engineering organizations control over the intelligence powering the AI tools they already use. Instead of being locked into a single provider’s pricing, models, and roadmap, you can choose the best model for every task, manage spend centrally, and continuously measure quality as new models emerge.
Ready to see what Nexus can do for your engineering team? Request a demo and we’ll show you how it works on your own workloads.
Book a demo