跳到主要內容
AI News HubLIVE
來源內容 · 翻譯待補全3 分鐘閱讀

待翻譯:Unlocking Data Portability: Preventing Catalog Lock-in with REGISTER and UNREGISTER APIs

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:As organizations embrace the lakehouse architecture, data is shifting from proprietary...

待翻譯:Unlocking Data Portability: Preventing Catalog Lock-in with REGISTER and UNREGISTER APIs
回報錯誤

更正管道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Unlocking Data Portability: Preventing Catalog Lock-in with REGISTER and UNREGISTER APIs | Databricks Blog Skip to main content REGISTER and UNREGISTER provide a standard method to safely transfer a table's catalog management without copying data. Open table formats are now easily transferable between catalogs when stored in customer-owned buckets. UNREGISTER eliminates the risk of dangerous "split-brain" scenarios by ensuring the original catalog explicitly relinquishes control before a transfer As organizations embrace the lakehouse architecture, data is shifting from proprietary data warehouses into open storage and table formats, where it can be accessed across multiple engines, like Spark and Trino, without duplication. However, open formats are only one part of the openness equation. True openness also requires interoperability and flexibility in how data is managed and governed. For a lakehouse to fully deliver on its promise, organizations need the freedom to choose and move between catalogs as their architecture evolves. To support this portability, the REGISTER and UNREGISTER endpoints enable users to hand off a table between catalogs without rewriting, exporting, or copying a single file. REGISTER attaches an existing table to any IRC catalog. When moving a table, you can't just DROP it from the old catalog because that cleans up its underlying data and metadata. We added the UNREGISTER endpoint to the Apache Iceberg™ REST catalog specification so that you can tell the old catalog to forget about the table and return the exact pointer the next catalog needs to take over. In this post, we will take a closer look at how the open table format ecosystem is evolving, the key challenges these additions address, and how REGISTER, and the new UNREGISTER command, actually work. Background: The Catalog's Role When an engine queries a table, it first asks the catalog to load the table to ensure it has the latest state. The catalog returns the table's current metadata, including the location of the table’s data in object storage. From there, the engine uses that metadata to find the schema and reads the Parquet straight from object storage. This coordination is essential for open table formats because all the important metadata - schema, history, statistics - and the data itself lives in storage, completely decoupled from compute. By acting as the central authority for that latest state, the catalog coordinates commits and ensures that two writers can never silently overwrite each other. Figure 1: The Read Path The engine asks the catalog to load the table. The catalog returns the table's current metadata and its location in object storage. The engine reads the metadata and the data files directly from your bucket. How the REGISTER endpoint works Attaching an existing table to a catalog is simply a matter of handing over the metadata location from loading the table. This is exactly what REGISTER does through the endpoint already defined in the Iceberg REST specification. Figure 2: The REGISTER Operation A client sends a POST request with the desired table name and the URI of its existing metadata location. After basic validation, the catalog writes a single record linking the table name to the existing metadata.json. No data files are copied, moved, or altered. The catch is that REGISTER alone would create a “split-brain” scenario, where two catalogs think they own the table and will coordinate commits. Because Iceberg catalogs are independent and do not communicate, REGISTER alone adds an entry to the new catalog while leaving the old one completely active. If two catalogs both believe they are the sole owner of the table, neither will throw an error, but the first write operation will fork the table, leading to inconsistent query results and silent data loss. Figure 3: The Split-Brain Scenario A write made through Catalog A will be completely invisible to Catalog B, permanently destroying your single source of truth. The missing piece was UNREGISTER Historically, the ecosystem lacked a standard way to cleanly terminate a catalog's management of a table. Running DROP TABLE doesn’t work because it deletes the table’s data and metadata! Without UNREGISTER, the hand-off equation for REGISTER was missing its second half. We contributed UNREGISTER to the Apache Iceberg™ REST catalog specification to remove the table's entry from the managing catalog without touching a single underlying data file. Crucially, it returns the table's latest metadata location - the exact pointer the next catalog needs to assume commit coordination. By ensuring the original catalog explicitly relinquishes control, this eliminates the risk of a split-brain scenario. The request is an empty POST to the table resource: POST /v1/{prefix}/namespaces/{namespace}/tables/{table}/unregister The response is the pointer the next catalog needs: Figure 4: The handoff Process UNREGISTER removes the table's entry from Catalog A (files stay exactly where they are). The response returns the table's latest metadata location. Catalog B picks up the pointer using the REGISTER endpoint. Catalog B is now the table's sole managing catalog. By combining these three steps - unregister, receive the location, and register - the handoff ensures the table has exactly one managing catalog at all times. (Note: A production migration still involves operational steps, like safely stopping writers and repointing jobs, which we will cover in a follow-up post.) Get started With REGISTER and UNREGISTER, you now have the freedom to move your tables to other catalogs. We’re adding this capability because open-source portability is critical for our customers. Unity Catalog remains the most open lakehouse for managing your data, providing the unified governance the agentic era requires. It provides the context layer for your ontology, delivers best-in-class access control and observability across data and agents, and delivers flexibility across clouds, regions, and compute. To try out REGISTER and UNREGISTER on Unity Catalog in private preview, reach out to your account team. Get the latest posts in your inbox Subscribe to our blog and get the latest posts delivered to your inbox. Sign up View all blogs

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級後設資料。
  • As organizations embrace the lakehouse architecture, data is shifting from proprietary...

技術影響

可能影響 Agent 架構、工具呼叫、工作流自動化和產品整合。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。