跳到主要內容
AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:From Data to Dialogue: How S&P Global Energy Made Its Structured Data Estate Conversational with Databricks Genie Agents and MCP

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:S&P Global's goal was to fundamentally improve how customers discover and consume...

待翻譯:From Data to Dialogue: How S&P Global Energy Made Its Structured Data Estate Conversational with Databricks Genie Agents and MCP
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

From Data to Dialogue: How S&P Global Energy Made Its Structured Data Estate Conversational with Databricks Genie Agents and MCP | Databricks Blog Skip to main content • Domain experts curate focused Genie Agents per dataset group without writing agent code, establishing a governed semantic layer. • Genie Agents act as managed MCP servers that are composed via a FastMCP proxy into composite endpoints for cross-domain queries. • This architecture significantly shortened time-to-market for conversational data products while preserving Unity Catalog governance. S&P Global's goal was to fundamentally improve how customers discover and consume insights across our data and research products. While AI powered search and summarization are important, the larger business value comes from enabling faster decision making through natural language access to trusted data, richer cross-commodity analytics, and the ability to connect insights that traditionally exist in separate business lines. AI agents help customers uncover relationships, generate research more efficiently, and derive actionable intelligence from a broader set of information than was previously possible. —Priyanka John, Vice President, S&P Global Energy If you have ever tried to make a large, complex structured data estate available to AI agents and assistants, you have likely run into the same wall we did: agents are only as good as the context they can reach, and enterprise data rarely lives in one neat, well-documented place. At S&P Global Energy, our data spans Chemicals, Crude Oil, Refined Products, Gas & Power, Liquified Natural Gas (LNG) and more — and each commodity is itself a rich family of datasets. LNG alone includes facility specifications, cargos , outages, supply and demand fundamentals, netbacks, historical and forecast prices, and contracts. Chemicals span capacity, production, utilization, trade, demand by end use and by derivative, inventory change, and country- and region-level supply–demand balances. Our other commodities follow similar patterns. This data lives across Databricks and several non-Databricks sources. Our goal was ambitious but simple to state: make our entire structured data estate available for external consumption by AI agents through the Model Context Protocol (MCP) — so that our customers’ agents and assistants, as well as our own, could ask questions in natural language and get trusted, governed answers. We evaluated several approaches. What worked best for us, by a wide margin, was Databricks Genie Agents, exposed as managed MCP servers, composed into domain-specific bundles with an MCP proxy layer. In this post, you’ll learn: How our subject matter experts (SMEs) curate focused Genie Agents — one per dataset group within each commodity — without writing a single line of agent code. How each Genie Agent automatically becomes a governed MCP server, ready to plug into any MCP-compatible client or agent. How we use a FastMCP-based proxy to compose multiple Genie MCP servers into composite endpoints for cross-domain questions. Why this architecture dramatically shortened our time to market for AI-powered data products. Genie Agents let our domain experts productize their knowledge of the data directly. What used to take a full development cycle now takes days, and every answer stays inside our governance boundary.—Priyanka John, Vice President, S&P Global Energy The challenge: Structured data is easy to store, hard to converse with Large language models are remarkably good at conversation and reasoning, but they cannot answer questions about your data unless you build a bridge to it. For structured enterprise data, that bridge has historically meant one of the following: Hand-built text-to-SQL pipelines — powerful, but brittle. Every schema change, every ambiguous column name, every domain-specific metric definition (“What counts as an outage day?”) becomes an engineering task. Custom APIs per use case — each new question pattern needs a new endpoint, a new sprint, a new release. Exporting data into external AI tools — which duplicates data, breaks freshness, and steps outside your governance perimeter. Each of these approaches shares the same problem: the people who understand the data best — our SMEs and analysts — are not the people building the access layer. Every insight had to pass through an engineering backlog. Our time to market for a new conversational data experience was measured in months. We needed an approach where domain experts could curate and publish conversational access to data directly, engineering could standardize how agents connect, and governance stayed centralized. That is exactly what Genie Agents plus MCP gave us. The architecture: Genie Agents as the semantic layer, MCP as the contract Our architecture has three layers, and each layer is owned by the people best suited to it. The diagram above shows the end-to-end flow — including what sits inside the S&P Global Energy network and what sits in the external client environment. Layer 1: SMEs curate one Genie Agent per dataset group This is where the magic starts, and notably, it requires no code. Our SMEs begin by selecting the tables relevant to a business domain: If the tables already live in Databricks, they use them directly through Unity Catalog. If the data lives in a non-Databricks source, they bring it in through Lakehouse Federation connectors — no data movement, no duplicate pipelines. The federated tables appear alongside native tables and inherit the same governance. They then group related tables and create one Genie Agent per dataset group — not one giant agent per commodity. Each sub-category of a commodity becomes its own focused Genie Agent. Within LNG, for example: LNG Assets & Contracts Genie Agent — assets, operators, capacity forecasts, and long-term contracts LNG Cargo Genie Agent — cargo tracking, fixtures, origins/destinations, and commercial terms LNG Tenders Genie Agent — tenders with issuers, volumes, and delivery windows LNG Outages Genie Agent — outages and maintenance events with capacity impact LNG Supply & Demand Genie Agent — fundamentals with regional splits and scenario history LNG Netbacks Genie Agent — netbacks from hub prices, freight, boil-off, and losses LNG Prices Genie Agent — historical and forecast price curves Every other commodity follows the same pattern with its own sub-categories. Chemicals, for instance, has group-level Genie Agents for capacity, production, capacity utilization, trade, demand by end use and by derivative, inventory change, and country- and region-level supply–demand balances; Crude Oil, Refined Products, and Gas & Power are organized similarly. The result is a fleet of small, sharply scoped Genie Agents rather than a handful of sprawling disconnected AI tools. Inside each agent, SMEs add the context that makes text-to-SQL actually work in the real world: descriptions of tables and columns, example queries, trusted assets for high-stakes metrics, and business definitions (for example “floating storage is defined as cargoes idling for 3 days or more in vessels travelling below a threshold speed”). This is the step that generic text-to-SQL solutions skip — and it is the step that determines whether users trust the answers. The key organizational insight: curation became a domain activity, not an engineering activity. The person who knows what “floating storage” means in an LNG context is the person teaching a Genie what it means. Layer 2: Every Genie Agent is automatically an MCP server Here is where Databricks did the heavy lifting for us. Each Genie Agent is exposed as a Databricks managed MCP server out of the box, at an endpoint of the form: https:///api/2.0/mcp/genie/{genie_space_id} There is nothing to deploy and nothing to host. Each server exposes a small, clean tool surface — essentially two tools per agent: A query tool (genie_query_space) — the agent submits a natural language question to the agent. A response tool (genie_poll_response) — the agent polls with the same conversation and message ID to retrieve the full response once it's ready, including the generated SQL and result set. This two-tool, ask-then-poll pattern turns out to be a great fit for agentic workloads: questions run asynchronously against a SQL warehouse, and the agent polls with the conversation and message ID returned by the query tool until the response is ready. Just as importantly, these managed servers are governed by Unity Catalog. A Genie Agent — or the user behind it — can only reach the agents and underlying tables they have permission to see. Authentication is handled by the platform. We did not have to build a security layer around our AI access; we inherited the one we already had. Layer 3: Composing group Genie Agents into commodity bundles with a FastMCP proxy One Genie Agent per dataset group keeps each agent focused and accurate. But real business questions routinely cross groups: “How did the recent outages at Sabine Pass affect cargo premiums into Asia?” touches both the Outages and Cargo Genie Agents at once, and cross-commodity questions like “How are naphtha prices affecting chemical production margins?” reach across Refined Products and Chemicals Genie Agents. Rather than building one giant agent (which degrades answer quality) or forcing every client to configure a dozen separate servers, we used FastMCP’s proxy and composition capabilities to create composite MCP endpoints — typically one per commodity, mounting that commodity’s group-level Genie MCP servers behind a single server with name-spaced tools. Higher-level composites can bundle several commodities the same way: from fastmcp import FastMCP # Each group-level Genie Agent is a managed MCP server on Databricks cargo = FastMCP.as_proxy(genie_mcp_config(“lng_cargo_agent_id”), name=“cargo”) outages = FastMCP.as_proxy(genie_mcp_config(“lng_outages_agent_id”), name=“outages”) netbacks = FastMCP.as_proxy(genie_mcp_config(“lng_netbacks_agent_id”), name=“netbacks”) # Compose the group Genies into one commodity bundle lng = FastMCP(name=“lng-composite”) lng.mount(cargo, prefix=“cargo”) lng.mount(outages, prefix=“outages”) lng.mount(netbacks, prefix=“netbacks”) # The same pattern repeats for Chemicals, Crude Oil, Refined Products, Coal … (Illustrative snippet — adapt to your FastMCP version and auth setup.) The result: an agent connects to one composite endpoint per commodity and sees a curated set of group tools — cargo_genie_query_agent, outages_genie_query_agent, netbacks_genie_query_agent, and so on, each paired with its genie_poll_response counterpart. The agent’s LLM decides which group Genie to route a question to, or fans a cross-group question out across several, then synthesizes the results. This gave us the best of both worlds: narrow, high-accuracy group-level Genie Agents underneath, and broad, commodity- and estate-wide conversational access on top. What changed for the business For our business stakeholders, the technical details above translate into a few very tangible outcomes. Time to market collapsed. Previously, standing up a new conversational data experience meant a full development cycle: requirements, API design, text-to-SQL engineering, testing, deployment. With this architecture, launching a new dataset group — or an entire commodity — means an SME creates and curates the corresponding Genie Agents — the MCP endpoint exists the moment the agent does. SMEs became publishers, not requesters. The domain experts who understand LNG cargoes or chemicals supply–demand balances no longer file tickets to get their data exposed; they curate a Genie Agent and it is live. Engineering effort shifted from building bespoke access layers to maintaining one thin, reusable proxy layer. Governance came built-in. Every question an agent asks runs through Unity Catalog permiss [truncated for AI cost control]

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • S&P Global's goal was to fundamentally improve how customers discover and consume...

技術影響

可能影響 Agent 架構、工具調用、工作流自動化和產品集成。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。