跳到主要內容
AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:Using AI to chart a course for our post-quantum migration

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We’re building CryptoLabe, an internal AI-powered tool that discovers cryptography across our codebase, surfaces dependencies, and helps us progress toward a full post-quantum migration by 2029. Here’s what we’ve learned so far.

來源Cloudflare AI Blog作者: Sharon Goldberg
待翻譯:Using AI to chart a course for our post-quantum migration
回報錯誤

更正管道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

As laboratories around the world race to build out a cryptographically relevant quantum computer, we at Cloudflare are racing towards a 2029 target deadline for full post-quantum readiness. While we’ve already transitioned many of our products to post-quantum encryption, we still have work to do to support post-quantum authentication and achieve full post-quantum readiness across our platform. We’re taking a maximalist stance (“PQ everything!”), because as an infrastructure provider to the world, we want to give our customers the peace of mind that using Cloudflare ensures that their traffic is future-proofed against quantum adversaries. But how does one accomplish such a massive migration at an organization of our size and scale? After all, cryptography is the base layer for almost all of the world’s digital systems, including the software services and the networking protocols that power our platform. To drive our PQ migration, we have three key goals. First, we want to help our product and engineering teams understand how cryptography is being used and how they should be upgrading it. This should cover both the upgrades to post-quantum encryption and to post-quantum authentication. Many of our products have already been upgraded to post-quantum encryption over TLS 1.3, but we still want to cover the long tail of TLS connections, as well as upgrade any other uses of public-key encryption. Meanwhile, it’s still early days for our deployment of post-quantum authentication. Next, we want to provide progress metrics for the migration. These might include per-repository and per-product counts of the use of classical and post-quantum cryptography. Finally, we want to surface prerequisites early. If our products or platform rely on protocols that don’t yet have a PQ migration plan (because PQ variants of the system have not yet been considered, because PQ standards do not exist or lack consensus, or because software libraries or other key ecosystem components do not yet have PQ support), then we need to know now. That way we can work with the relevant stakeholders, standards bodies and ecosystems to help drive their PQ migration plans, so that we can meet our own 2029 PQ migration timeline. This post is the story of how we’re going about this. We explain how we turned to AI to help us solve some of our problems and how we’re developing an internal tool called CryptoLabe to help us. CryptoLabe is named after the mariner’s astrolabe, a navigation instrument refined by Portuguese navigators. Just as an astrolabe helped sailors determine where they were and chart a course, CryptoLabe helps us discover cryptography in our code, understand how it is used, and chart a path to post-quantum migration. CryptoLabe is highly specialized to our internal systems (our repositories, our ticketing systems, and internal documentation processes) and still evolving as we continue its development, so we aren’t making it available to customers. Nevertheless, we are sharing our learnings so that other organizations can build upon our efforts as they work through their own PQ migration journey. The scale of the problem The software that powers most Cloudflare products lives inside our single centralized source control management platform. This means we can find most uses of cryptography across our platform by just looking through our codebase. While the centralization of our codebase is a marked advantage for us, we still need to contend with three challenges that come with the scale of this problem. First, our code is spread across many repositories. Second, cryptography rarely announces itself plainly in the code. Instead, it hides in shared libraries that a repository imports but may or may not actually call upstream and protocol defaults, like a TLS 1.3 listener that is configured to negotiate a classical key exchange such as X25519 rather than post-quantum X25519MLKEM768 configuration files that select algorithms far away from the code that uses them, like a TLS responder whose key exchange protocols are pinned in a YAML file stored in a different repository code paths that are dead, test-only, or on a path to being deprecated Third, cryptography discovery is about more than just pattern matching. Grepping for certain algorithm names (e.g. “RSA” or “X25519”) overcounts, because it finds cryptography in unused code. Grepping also undercounts, because it misses defaults and indirect uses in dependencies and configuration. Most importantly, it can't tell you how the cryptography is used. A classical ECDSA signature could be part of a JWT, IPsec, TLS, or SSH, and each has a completely different migration path. Many uses also depend on the other side of the connection: a TLS server may support both post-quantum key exchange and classical key exchange; the one it chooses to use would depend on the client. Turning to AI It turns out that AI is pretty good at doing more than just grepping. A model can search a codebase, follow evidence across files, and return structured analysis. It can also enrich findings by pulling information from other sources, like our internal documentation and ticketing systems. In fact, AI can even explain how cryptography is being used and how it should be updated. We’ve been putting that idea to the test as we develop CryptoLabe. As we said before, our first two goals are to (1) discover and understand the use of cryptography in our codebase, and also (2) to get metrics on the state of our PQ migration. Towards these goals, our current implementation of CryptoLabe performs scans in two stages, as shown in the figure below. The first “discovery” stage starts by mapping the repository. It then searches for cryptography through source, configuration, manifests, lockfiles, scripts, tests, and documentation. Among other things, the scan looks for the use of cryptography like key agreement, signatures, asymmetric encryption, PKI, tokens, credentials, hardware security module integrations, and more. This discovery stage produces a set of "raw observations." Each raw observation feeds a run of the second stage. This “analysis” stage first re-checks the observation against the source code. It then investigates how the cryptographic operation is used at runtime, what role the repository plays, and which internal or external parties it depends on. When necessary, it can inspect related code in other repositories to complete the analysis. Finally, it takes a pass over its own conclusions, searching for missing or conflicting evidence such as configuration overrides, test-only code, or incorrect assumptions about runtime behavior. Next, the model assigns a classification to the finding. If there is not enough evidence to assign a classification, the model assigns More evidence needed, External dependency, or Unknown rather than guessing. This is the current list of classifications used by CryptoLabe, containing catch-all classifiers which will likely be refined as we proceed through our migration. (As an example, we could refine our classifiers by splitting the “encryption” classifier into key agreement and HPKE; you get the idea.) Classification Examples Classical encryption This is a catch-all category that finds cases of elliptic-curve Diffie-Hellman key exchange (ECDHE) (e.g., X25519, P-256, P-384), RSA key agreement or other uses of public-key encryption (e.g., HPKE). These are broken by a quantum computer running Shor's algorithm, which puts them at risk of harvest-now-decrypt-later attacks. Classical signature This is a catch-all category that finds use of an RSA signature or elliptic-curve (ECDSA) signature in anything, for example a certificate, a TLS handshake, another protocol handshake. These signatures are broken by Shor's algorithm. Classical token We found a lot of RS256 or ES256 JWT tokens, so we created a special classification for them. These are JWTs that use classical RSA and ECDSA signatures; RFC 9964 defines a post-quantum replacement using ML-DSA. PQ-ready hybrid key exchange Finds hybrid post-quantum key exchange in TLS 1.3, i.e. X25519MLKEM768. This is the most prevalent use of PQ encryption in our codebase. PQ-ready Finds other uses of post-quantum cryptography that are not X25519MLKEM768 in TLS 1.3, like ML-DSA. Finally, it generates a report that serves two audiences: (1) product managers who need to understand what the migration means for their product, and (2) engineers that need enough detail to execute the migration. Here’s a (cropped) view of one of our reports: While we’ve been iteratively reviewing findings against the source code and with relevant engineers, we do not yet have a ground-truth dataset for reproducibly comparing different versions of the prompts we’ve tried for CryptoLabe. Built on Cloudflare’s Developer Platform We built CryptoLabe on Cloudflare's Developer Platform. Here’s the architecture: CryptoLabe runs across two Cloudflare Workers. There’s a scanner Worker that runs the scans. And there’s an inventory Worker that serves the dashboard, exposes the API, and stores everything in a D1 database. The two communicate through Service Bindings. A scan starts when someone requests it from the dashboard, and the inventory Worker passes the request to the scanner. Orchestrating a scan We need a way to keep a scan alive and on track from start to finish, without building our own job orchestration system. We did this with Agents SDK. Each repository gets its own persistent coordinator built on a Durable Object (DO). A bounded queue in front of the coordinators limits how many scans run at once. When a scan's turn comes, the coordinator tracks its progress and handles cancellation, retries, and recovery. The coordinator doesn't do the analysis itself. It hands the work to Cloudflare Workflows, so that they can persist progress and automatically retry failed steps. The coordinator moves each repository through four stages: discovery Workflow (the first scanning stage that produces raw observations) deep analysis Workflow (the second stage, run on each raw observation) merge Workflow (that builds a list of findings for a given repository, including combining repeated or similar finds) publish workflow (that hands results back to the inventory Worker) The first two workflows need the model to have access to the repository's code. We want this access to be isolated, so we don’t risk damaging the codebase. That’s why CryptoLabe downloads the repository once, at an exact commit, at the start of each scan, and then stores that snapshot in R2. Each Workflow then restores the snapshot into a fresh, short-lived Cloudflare Sandbox, an isolated container. The model then works with the Sandbox through a small set of read-only tools on an immutable snapshot of the code, even if the codebase changes while the scan is still running. Calling the model at scale If we want to scan through all of our (many!) repositories, we have to worry about both cost and capacity. For cost, the model loop sends its requests through AI Gateway to cost-effective open-weight models hosted on Workers AI. Putting the model behind AI Gateway also makes it easy to switch models as better or cheaper ones become available. Capacity became a problem once we scanned many repositories at once. Bursts of model requests began triggering HTTP 429 (rate limit) responses from AI Gateway, and scans retrying independently only made the bursts worse. We solved this with a single, global Durable Object that paces every model request across all scans, including retries. When any scan hits a rate limit, the cooldown is shared and all scans back off together, so concurrent scans share the available capacity instead of competing for it. Prerequisites and hard cases Let’s now get into our third goal: surfacing prerequisites and hard cases early. A lot of ink has been spilled about ecosystem readiness for the PQ [truncated for AI cost control]

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級後設資料。
  • We’re building CryptoLabe, an internal AI-powered tool that discovers cryptography across our codebase, surfaces dependencies, and helps us progress toward a full post-quantum mig…

技術影響

可能影響 Agent 架構、工具呼叫、工作流自動化和產品整合。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。