Skip to content
AI News HubLIVE
Original source5 min read

How Databricks rolls out frontier models to 14,000 employees on Day 1

Summary

Providing our employees access to frontier AI capabilities is a top priority at Databricks, and consequently...

How Databricks rolls out frontier models to 14,000 employees on Day 1
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

How Databricks rolls out frontier models to 14,000 employees on Day 1 | Databricks Blog

Skip to main content

Providing our employees access to frontier AI capabilities is a top priority at Databricks, and consequently, it is important to us for them to use new models instantly when they become available. At the same time, it is nontrivial to give more than 10,000 people rapid access to a new model because:

Models that are marketed as frontier often aren’t. For example, Opus 5.0 was more expensive and ranked lower on both quantitative and qualitative quality scores among our engineers compared with Opus 4.8. Migrating to a model that regresses the frontier can meaningfully hurt a company, rather than help it. In our experience, substantial care is required when evaluating models before migrating workloads en masse to new models.

Naive use of a new model can explode costs. When we released GPT Astra to a control group with no associated cost mitigations, the average developer spent 60% more than before gaining access to Astra. An overnight 60% cost increase with a user population of more than 10,000 is very difficult for a company to plan around. Once we better understood which tasks Astra is uniquely great at, we were able to steer usage to substantially reduce overall costs.

This post discusses a set of techniques we’ve employed to give most employees at Databricks “Day 1” access to new models, while allowing us to assess whether models are indeed good long-term workhorses. These techniques rely heavily on Unity Gateway to adaptively release, evaluate, and incorporate new models. The week of September 21 was a critical test of these abilities when Opus 5, GPT-6 Sol and GPT-Luna were released in rapid succession. During the week, Databricks provided all employees with Day 1 access, and by Day 3, we had gathered enough data to confirm that these models were on the efficiency frontier, leading to their incorporation into our broader infrastructure.

The Model Release Lifecycle

At a high level, new model releases at Databricks go through a pipeline that looks as follows:

Immediately make new models available to all employees, on an "experimental" basis.

Constrain usage of new models based on a per-user budget.

After gathering enough data, decide whether to promote the model to production (or even make it the default).

Step 1: Make new models available immediately

To make model management across closed and open model providers easier, we leverage our own Databricks Unity Gateway for all internal use. This is our central hub for AI governance, cost management, and observability, so it's natural that we start here.

The Gateway is where we enable all employees to access the newly released model. Server-side configuration is not enough, though. Our employees are using Claude Code, Codex, and the Omnigent meta-harness on their laptops, and we need to distribute the new model configuration to them.

That's where Unity Gateway CLI (UG CLI) comes in. The UG CLI is already running on everyone's laptop, deployed via our Mobile Device Management. Whenever someone starts Claude Code, Codex, or Omnigent, the UG CLI runs to check for new models, tools and skills, and updates the local harness's configuration. UG also lets us centrally designate default vs. experimental models, prepare models for smart routing, and collect traces to evaluate each model rollout.

We configured Unity Gateway to push experimental configurations for Opus 5.5 and Sol 6. These models now show up with this tag, so that employees can select it but understand it's a new model that may or may not be best in class or stick around forever:

Claude Code / model output clearly designates Ous 5.5 as Experimental

Step 2: Constrain usage using a per-user budget

We have previously written about how we configure per-user budgets for AI spend. Since then, we have expanded our total budget architecture to include four principal budgets, each defined on a per-user basis:

Monthly maximum: Each user has an overall monthly spending ceiling across all models.

Daily runaway limit: Every user has a daily max, which can be raised directly in Slack to avoid accidental spending from a runaway session.

[New!] Quality frontier budget: We allocate a certain fraction of the monthly budget to the most premium models at the quality frontier, such as GPT Astra and Claude Fable. (Fable is not currently rolled out internally due to Anthropic data retention policies, but we are working closely to implement their new policy.) This reflects the intent that these models should not be used as daily drivers, but instead selected for specialized tasks where they are uniquely suited, to justify the 2-3x cost increase over the next quality tier.

[New!] Experimental budget: Another fraction of the monthly budget is allocated for the usage of new, untested models. Here, we aim to balance speed of adoption with the downside risk of widely exposing a model that is not on the efficiency frontier.

On day one of the model launch, we made Opus 5.5 and Sol 6 available to all employees via Unity Gateway and tagged them for the experimental budget. We then used the next few days to collect data to decide what to do next — drop the experimental tag or remove it from the model catalog that our developers see.

Budget setup overview, reflecting the four budgets

Step 3: Promote or drop the model

In order to determine whether the model is at the efficiency frontier, we rely on three signals:

Benchmark data: We have a set of private benchmarks that test a suite of tasks, including offline benchmarks such as document reasoning, workspace search, and our own Genie product, as well as online benchmarks in which we run two models side-by-side and compare outputs for pull request creation. We continue to expand and tune these benchmarks; in an ideal world, our benchmarks are sufficient to quickly determine the cost and quality of any new model release.

User-reported quality: The experimental model release provides a wealth of anecdotal data on how people feel about the new model. We find that a group of power users is eager to try out new models and compare their experiences over Slack and survey responses.

Cost tracking via OpenTelemetry traces: Unity Gateway logs all traces in a central location, along with cost information. We can compare how pilot users spent money on the prior generation of models versus the latest models on a per-session basis.  This does not necessarily tell us quality, but it gives us a good measure of the cost.

For Opus 5.5 and Sol 6, all three metrics tell us a pretty consistent story.

Benchmarks, such as our OfficeQA Pro V2, show that Opus 5.5 is clearly on the cost/quality frontier, a huge step up on both axes from Opus 5. GPT-6 Sol scores somewhere in between GPT-5.6 Sol and GPT-5.6 Terra on both cost and quality.

User reports broadly agree that for engineering and debugging tasks (the vast majority of our early adopters are engineers), Opus 5.5 is a big step up in quality from Opus 5 and Opus 4.8, and its writing style is vastly preferred. Whereas GPT-6 Sol is occasionally a downgrade in quality compared to GPT-5.6 Sol.

Cost tracking allowed us to compare early adopters' usage with the same group's usage a week earlier. Maintaining the same cohort proved crucial because early adopters tend to be power users of AI rather than average users.

We wanted to normalize costs to a $/session basis, since users who are trying out a new model will sometimes increase their usage in terms of number of sessions as they experiment. We found that using a basic $/session comparison was still misleading because the distribution of sessions was also changing: early adopters were trying harder problems with the new models than their average session.

As a result, we stratified sessions based on whether they were single- or multi-turn and whether they made any file edits, and then reweighted the distribution accordingly. The table below shows the results for Opus 5.5 vs. Opus 4.8 and GPT-6 Sol vs. GPT-5.6 Sol.

Cost comparison

Old Model (avg $/session)

New Model (avg $/session)

Delta

Opus 4.8 vs. Opus 5.5

$5.94/session (Opus 4.8)

$4.23 (Opus 5.5)

−29%

GPT-5.6 Sol vs. GPT-6 Sol

$4.52/session (GPT-5.6 Sol)

$2.34/session (GPT-6 Sol)

−48%

The GPT numbers are not too surprising given that the price was cut by 50%. But we were pleased that Opus 5.5 also represents a significant price reduction for our real workloads, given our prior experience with Opus 5.

What we decided

We were able to provide employees experimental access to Opus 5 and GPT-6 Sol and Luna on Day 1 of model launch. Within three days, we had gathered enough data to confirm that these models were on the efficiency frontier and decided to move them out of the experimental budget and into standard circulation as generally available models.

Over the next week, we will go a step further for Opus 5.5 to make it the default for Claude Code, given its clear position as higher-quality and lower-cost than its predecessors.

Our experience with GPT-6 Sol suggests that it will not replace GPT-5.6 Sol as the default for Codex. However, we will include GPT-6 Sol in our smart router’s toolkit, given its cost advantage over 5.6 Sol.

Overall, we found this playbook effective at quickly assessing model quality and cost, enabling us to rapidly adopt the latest models that prove their worth. This flexibility matters now more than ever, with new models arriving almost daily.

Get the latest posts in your inbox

Subscribe to our blog and get the latest posts delivered to your inbox.

Sign up

View all blogs

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Providing our employees access to frontier AI capabilities is a top priority at Databricks, and consequently...

Highlights and analysis are generated automatically and may contain errors. Check the original source.