本文にスキップ
AI News HubLIVE
原典の内容 · 翻訳・分析待ち3 分で読了

翻訳待ち:Unlocking Data Portability: Preventing Catalog Lock-in with REGISTER and UNREGISTER APIs

記事の要約

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:As organizations embrace the lakehouse architecture, data is shifting from proprietary...

翻訳待ち:Unlocking Data Portability: Preventing Catalog Lock-in with REGISTER and UNREGISTER APIs
誤りを報告

訂正窓口はまだ利用できません。記事情報をコピーして保存できます。

訂正案内
本文へ

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Unlocking Data Portability: Preventing Catalog Lock-in with REGISTER and UNREGISTER APIs | Databricks Blog Skip to main content REGISTER and UNREGISTER provide a standard method to safely transfer a table's catalog management without copying data. Open table formats are now easily transferable between catalogs when stored in customer-owned buckets. UNREGISTER eliminates the risk of dangerous "split-brain" scenarios by ensuring the original catalog explicitly relinquishes control before a transfer As organizations embrace the lakehouse architecture, data is shifting from proprietary data warehouses into open storage and table formats, where it can be accessed across multiple engines, like Spark and Trino, without duplication. However, open formats are only one part of the openness equation. True openness also requires interoperability and flexibility in how data is managed and governed. For a lakehouse to fully deliver on its promise, organizations need the freedom to choose and move between catalogs as their architecture evolves. To support this portability, the REGISTER and UNREGISTER endpoints enable users to hand off a table between catalogs without rewriting, exporting, or copying a single file. REGISTER attaches an existing table to any IRC catalog. When moving a table, you can't just DROP it from the old catalog because that cleans up its underlying data and metadata. We added the UNREGISTER endpoint to the Apache Iceberg™ REST catalog specification so that you can tell the old catalog to forget about the table and return the exact pointer the next catalog needs to take over. In this post, we will take a closer look at how the open table format ecosystem is evolving, the key challenges these additions address, and how REGISTER, and the new UNREGISTER command, actually work. Background: The Catalog's Role When an engine queries a table, it first asks the catalog to load the table to ensure it has the latest state. The catalog returns the table's current metadata, including the location of the table’s data in object storage. From there, the engine uses that metadata to find the schema and reads the Parquet straight from object storage. This coordination is essential for open table formats because all the important metadata - schema, history, statistics - and the data itself lives in storage, completely decoupled from compute. By acting as the central authority for that latest state, the catalog coordinates commits and ensures that two writers can never silently overwrite each other. Figure 1: The Read Path The engine asks the catalog to load the table. The catalog returns the table's current metadata and its location in object storage. The engine reads the metadata and the data files directly from your bucket. How the REGISTER endpoint works Attaching an existing table to a catalog is simply a matter of handing over the metadata location from loading the table. This is exactly what REGISTER does through the endpoint already defined in the Iceberg REST specification. Figure 2: The REGISTER Operation A client sends a POST request with the desired table name and the URI of its existing metadata location. After basic validation, the catalog writes a single record linking the table name to the existing metadata.json. No data files are copied, moved, or altered. The catch is that REGISTER alone would create a “split-brain” scenario, where two catalogs think they own the table and will coordinate commits. Because Iceberg catalogs are independent and do not communicate, REGISTER alone adds an entry to the new catalog while leaving the old one completely active. If two catalogs both believe they are the sole owner of the table, neither will throw an error, but the first write operation will fork the table, leading to inconsistent query results and silent data loss. Figure 3: The Split-Brain Scenario A write made through Catalog A will be completely invisible to Catalog B, permanently destroying your single source of truth. The missing piece was UNREGISTER Historically, the ecosystem lacked a standard way to cleanly terminate a catalog's management of a table. Running DROP TABLE doesn’t work because it deletes the table’s data and metadata! Without UNREGISTER, the hand-off equation for REGISTER was missing its second half. We contributed UNREGISTER to the Apache Iceberg™ REST catalog specification to remove the table's entry from the managing catalog without touching a single underlying data file. Crucially, it returns the table's latest metadata location - the exact pointer the next catalog needs to assume commit coordination. By ensuring the original catalog explicitly relinquishes control, this eliminates the risk of a split-brain scenario. The request is an empty POST to the table resource: POST /v1/{prefix}/namespaces/{namespace}/tables/{table}/unregister The response is the pointer the next catalog needs: Figure 4: The handoff Process UNREGISTER removes the table's entry from Catalog A (files stay exactly where they are). The response returns the table's latest metadata location. Catalog B picks up the pointer using the REGISTER endpoint. Catalog B is now the table's sole managing catalog. By combining these three steps - unregister, receive the location, and register - the handoff ensures the table has exactly one managing catalog at all times. (Note: A production migration still involves operational steps, like safely stopping writers and repointing jobs, which we will cover in a follow-up post.) Get started With REGISTER and UNREGISTER, you now have the freedom to move your tables to other catalogs. We’re adding this capability because open-source portability is critical for our customers. Unity Catalog remains the most open lakehouse for managing your data, providing the unified governance the agentic era requires. It provides the context layer for your ontology, delivers best-in-class access control and observability across data and agents, and delivers flexibility across clouds, regions, and compute. To try out REGISTER and UNREGISTER on Unity Catalog in private preview, reach out to your account team. Get the latest posts in your inbox Subscribe to our blog and get the latest posts delivered to your inbox. Sign up View all blogs

要点と分析を開く

記事インテリジェンス

エンジニア上級

要点

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • As organizations embrace the lakehouse architecture, data is shifting from proprietary...

要点と分析は自動生成され、誤りを含む場合があります。原典をご確認ください。