AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。
Set Budgets and Alerts for Cloud Data Warehouse Costs | Databricks Blog Skip to main content Tag every SQL warehouse with team, cost centre, and use case at creation, then monitor attribution through system.billing.usage to bring unattributed spend to zero. Set layered budgets (per-team for accountability, account-wide as a safety net) and wire up pacing alerts that project month-end overruns days before they happen. Build cost dashboards on system tables and make the warehouse spend a metric every team owns, not one the platform team explains after the fact. Every analytics team has been here: a month-end invoice arrives, and a single SQL warehouse has blown through its budget, maybe a compute left running overnight, maybe an ad hoc one that auto-scaled to serve 80 Tableau users during a quarterly review. The cost is real, but the frustration runs deeper: nobody knew until it was too late. Proactive cost management flips that model. Instead of forensic bill analysis after the fact, you set guardrails before a dollar is spent - budgets are tied to warehouses, teams, or BI workloads. Alerts that fire at the thresholds you choose, and dashboards that surface usage trends in real time, all on the Databricks Data & AI Platform. If you're mid-migration from an on-premises system or another cloud data warehouse, now is the moment to put that governance in place. Why does reactive cost management fail? Data warehousing workloads present unique cost challenges. SQL warehouses serve interactive, human-driven workloads (BI dashboards, ad hoc analysis, embedded analytics) where demand is unpredictable and bursty. A single executive all-hands that triggers 200 simultaneous dashboard refreshes can spike costs in minutes. Most organizations start their FinOps journey the same way: someone downloads last month's usage CSV, opens a pivot table, and starts asking questions. This approach breaks down for four reasons: You can't see who's spending. Usage rolls up to the platform, not to the team or query that caused it, so that no one can be held to a number. No one owns a budget. Without spending targets tied to a warehouse or team, the analyst running SELECT * on a 2 TB table never feels the cost. The platform team quietly absorbs it. You find out too late. By the time last month's Excel reports surface the overspend, it's two billing cycles old and the decision that caused it is invisible. Manual review doesn't scale. Five warehouses may pass a monthly pivot-table review. However, once scaled to dozens, supporting hundreds of analysts across Tableau, Power BI, and Looker becomes unmanageable. Proactive cost management — attribution, budgets, alerts, dashboards, and optimization — addresses all five. Start with serverless: the first cost decision. Before setting up budgets, choose the right warehouse type. SQL Serverless warehouses are Databricks’ recommended option for most workloads. They start in seconds, scale down the moment queries finish, and let Intelligent Workload Management allocate compute for you, so you pay only for actual query execution, not idle readiness. If you're migrating from a legacy cloud data warehouse, this is a major shift. Most legacy platforms charge for always-on compute or require manual scaling decisions. Serverless removes that entire class of cost mistakes. Lumen Technologies saw this firsthand: Migrating two essential telecom systems from an on-premises data warehouse to SQL Serverless reduced compute costs by 30 -40% and increased query speed by 90%. The system now automatically scales to handle approximately 7 GB every 10 minutes, with Serverless removing the "always-on tax" typically associated with bursty, real-time workloads. The five pillars of proactive cost management The framework consists of five pillars. The first four make your spend visible; the fifth is what actually brings it down: Cost Attribution. Tag everything so you know who spent what on which warehouse. Budgets. Define monthly spending targets at the account, workspace, or team level. Alerts. Get notified the moment the spend approaches or exceeds a threshold. Dashboards. Visualize warehouse trends and make cost a first-class operational metric. Optimizations. Reduce what the work costs in the first place. Pillar 1: Cost attribution at two levels You can't manage what you can't attribute, and with warehouses that means two levels: the warehouse itself and the individual query. One tells you which warehouse spent the money; the other tells you who inside it did. Level 1: Tagging the warehouse SQL warehouses support custom tags as key:value pairs that propagate to system.billing.usage, linking every DBU consumed to a team, project, or cost centre. Tagging works the same way for every warehouse type — Classic, Pro, and Serverless are all tagged at the warehouse level, so you learn one pattern and apply it everywhere. Set tags in the UI under SQL Warehouses > your warehouse > Edit > Tags, or create the warehouse through the REST API: Tagging through the REST API Apply at least three dimensions: team (who owns it), cost center (who pays), and use case (BI serving, ad hoc, scheduled refresh). There isn't a native policy today that enforces warehouse tags, so anyone can create an untagged one. Build enforcement into your provisioning process instead — create warehouses with Terraform, Declarative Automation Bundles, or the REST API with tags in the definition, and treat the UI as the exception. The attribution-gap query later in this post catches whatever slips through. Level 2: Tagging the query Warehouse tags tell you a warehouse cost $4,000 last month. They don't tell you that 60% of it came from a single Power BI dashboard refreshing every 15 minutes. That gap matters because warehouses are shared — one BI warehouse serves dozens of dashboards, several dbt models, and a stream of ad hoc queries. Query tags (Public Preview) close it by attaching business context to individual SQL statements: The tags land in the query_tags column of system.query.history (Public Preview), alongside executed_by and statement_id, so you can group spend by dashboard, model, or cost centre rather than by warehouse. Some tools set these for you: starting with dbt-databricks 1.11.0, model queries are automatically tagged with the dbt model name, and Power BI passes workspace and dataset identifiers through the ADBC driver. Two caveats: query tags apply to SQL warehouse queries only, and because they live in query history rather than billing data, you join the two yourself to get per-query cost. Before writing any SQL, check the Cost page in Governance Hub (Beta) — an account-level view of spend, cost drivers, budgets, and tagging coverage, and the quickest way to see how much of your usage is attributable at all. An account admin enables it from the account console's Previews page. Pillar 2: Setting budgets in the account console Creating a budget In the Account Console, go to Usage > Budgets, click Add budget, and configure: Name: Descriptive (e.g., "Analytics SQL Warehouses, Monthly") Amount: Monthly target in USD Scope: Filter by workspaces and/or tags (e.g.,Team:Analytics) Email notifications: Recipients notified when spend reaches the budget amount Example budget configurations Budget Name Amount Scope (Tags) Alert Recipients BI Serving + Production $8,000/mo UseCase:BIServing Env:Prod [email protected] Analytics, Ad-Hoc $3,000/mo UseCase:AdHoc [email protected] Marketing Analytics $2,500/mo Team:Marketing [email protected] Account-Wide Safety Net $25,000/mo (all SQL usage) [email protected], [email protected] The key pattern is layered budgets: team-level budgets for accountability, plus an account-wide budget as a safety net that catches anything that slips through. Keep in mind that budgets are a monitoring mechanism, not a hard cap. They don't stop usage or prevent charges, so your bill can still exceed the amount. The goal is awareness and rapid response, not shutoffs that could break a production dashboard mid-refresh. GetYourGuide tested this directly: consolidating all their Looker workloads onto SQL Serverless lowered BI serving costs by ~20% and sped queries 35%, even though the team had expected Classic clusters to be cheaper. Warehouse-type choice is a budgeting decision worth re-validating, not a one-time call. Reading your burn rate Click any budget to see current spend vs. target, remaining budget, and a day-by-day burn-rate visualisation. For data warehousing, this chart is especially revealing: Linear daily burn means stable BI serving costs. Sawtooth with Monday/Friday spikes indicates ad hoc exploration clustering around workdays. A sudden step-up mid-month means a new dashboard was deployed without a cost review. Gradual upward drift means organic adoption is outpacing your budget assumptions. Pillar 3: Alerts, the nervous system of cost governance Budgets tell you where you stand. Alerts tell you when to act. Budget email alerts When you create a budget, add email recipients who are notified when spending exceeds the budget amount. Zero code is needed. For threshold-based alerting, create Databricks SQL Alerts that query system.billing.usage directly. Here are the patterns that matter most: Alert when the daily SQL warehouse spend exceeds a threshold Pacing alert: projected month-end spend will exceed budget This is the pattern teams get the most value from. Instead of alerting after the breach, it projects month-end spend and fires before that projection crosses your budget while you still have time to act: Schedule this every four hours and route to Slack or email. It gives teams days to react (right-size a warehouse, optimize an expensive query, defer a batch) before the budget is breached. Detect hidden costs from materialized view refreshes Materialized view refresh costs are the ones teams miss most often, because they bill as serverless pipeline usage rather than warehouse usage. A view set to refresh every 15 minutes against a source that updates twice a day spends the difference for no benefit. Match the refresh schedule to how often the underlying data actually changes, then watch the pipeline spend directly: Pillar 4: Dashboards and continuous monitoring Alerts are point-in-time triggers. Dashboards give continuous context. Databricks provides pre-built usage dashboards that account admins can import into any Unity Catalog-enabled workspace. These cover spend trends, top-N analysis, tag filtering, and per-warehouse drill-downs. For custom monitoring, two queries are particularly useful: SQL warehouse cost by team Unattributed SQL usage (attribution gap) Untagged SQL usage is your attribution gap, warehouse spend that can't be traced to any team. Drive it to zero. Raiffeisen Bank International built its own cost-monitoring and forecasting platform directly on Databricks system tables, with per-warehouse, per-user, and per-workload visibility. The payoff was cultural as much as technical: 3-4x faster workloads and, more importantly, when teams can see their own spend and are expected to explain it, behaviour changes without mandates. Pillar 5: Optimization, the pillar that actually changes the bill The first four pillars are about seeing your spending. This one is about lowering it, and it's the one most teams skip. Start with the two changes that improve cost and performance at the same time. Turn on predictive optimization for Unity Catalog managed tables. Databricks runs OPTIMIZE, VACUUM, and ANALYZE for you based on how each table is actually queried, and with CLUSTER BY AUTO it adjusts clustering keys as query patterns change. Less data scanned means faster queries and fewer DBUs, with no maintenance schedule to keep. It's on by default for accounts created on or after November 11, 2024, and reaching older accounts throug [truncated for AI cost control]