AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
Artificial intelligence systems are difficult to reproduce because their behavior depends on nondeterministic models and on data, configurations, policies, and external services that continue to change. But exact reproducibility is not always what operators, investigators, or auditors need. They often need something different: historical reconstructability. The central claim is that reconstructability is a system property established during execution, not inferred later from surviving artifacts. To make that property concrete, this article introduces Orrery, a reference architecture for preserving the runtime bindings and historical dependency states needed to reconstruct an AI system’s past execution. Imagine an AI-assisted decision that must be examined six months after it was made. Rerunning the system may produce a different answer, but an investigator may still need to determine which model endpoint, policy, retrieved data, tool contract, and runtime configuration were used at the time. That is the problem addressed here. The AI industry remains much better at evaluating capability than at preserving the knowledge required to understand past executions. Evaluation pipelines track benchmark scores, task success rates, retrieval precision and recall, tool-call success rates, latency, and cost, and they grow more elaborate with every release. In practice, this machinery often reduces to one acceptance question: What can this system do? That question captures capability, but not whether a past execution will remain understandable after the system changes. That missing property is historical reconstructability: the ability to establish which dependencies were used and under which conditions a particular execution occurred. It is related to, but distinct from, reproducibility. Reproducibility asks whether a result can be produced again under equivalent conditions. Reconstructability asks whether the conditions of the original execution can still be identified, even when repeating that execution would not produce an identical result. This distinction matters when a consequential AI execution completes successfully today and is challenged six months later. The model may still exist in a registry. The application revision may still exist in Git. The trace may still identify the request and the services it crossed. Yet none of these records necessarily reveals which model endpoint served the request, which input transformation and runtime configuration were applied, which data, feature state, or retrieved context entered the computation, or which policy state governed it. In an agentic system, the missing history may also include memory, tool contracts, delegation, and external effects. Even when a dependency identifier was recorded, the historical state to which it referred may no longer be available or resolvable. The problem is that an AI runtime may use dependencies whose historical identities and relationships are not preserved. This is an architectural problem, not merely a logging problem. The gap lies between what the runtime uses and what the surrounding infrastructure is required to preserve. I call this the preservation gap. The AI stack preserves models, code, deployments, and observations, but it does not require the effective historical configuration of a particular execution to remain available. The components may remain while the relations that made them one execution disappear. An AI runtime must assemble today’s execution; it is not necessarily required to remember exactly what it assembled yesterday. Reconstructability must therefore be established during execution, before later mutation erases the bindings that gave the execution its effective configuration. Hermetic builds illustrate the opposite design pattern. They close the dependency graph before execution begins by pinning inputs, versions, and external dependencies. AI runtimes face the reverse situation: Part of the dependency graph remains open until execution, as model endpoints may resolve through mutable aliases, data and feature state evolve, runtime configuration and policy change, and retrieved context is selected only when a request is processed. Agentic systems extend the graph further through memory, tools, delegation, and external effects. In a hermetic build, dependency closure is an input. In a modern AI runtime, part of it may be an output of the run. Version control records source evolution, transaction logs record state transitions, and distributed tracing records execution relationships. These mechanisms preserve different views of a run but not its materially relevant execution-specific configuration. A trace may retain execution flow, a deployment manifest the declared application state, and a model registry the model revision. What is often missing is the record of which artifacts and states were bound together in a particular execution. The dependency graph no longer closes at deployment A deployment records the configuration known before execution begins. AI systems are harder to reconstruct because materially relevant dependencies are selected during execution rather than fixed at deployment. At inference time, the system may resolve a model endpoint, input transformation, data or feature state, retrieved context, runtime configuration, safety controls, and policy. Agentic systems extend this composition through memory, tools, delegation, and external effects. These execution-specific selections are runtime bindings: facts about what the execution actually used. The following example makes this distinction concrete. Consider a hypothetical agentic execution identified as E891. During this execution, the system retrieves two documents, evaluates a policy, invokes a tool, and produces an external effect. The deployment may identify application revision 8f31c2 and model v17, while the execution itself binds prompt h31, retrieval index r42, entities d182@17 and d761@4, tool contract h42, policy v8, and ultimately effect e3. The deployment system could not have known this entire set in advance, because some of these dependencies were selected only as the execution proceeded. The difference between what a deployment declared and what an execution actually used becomes consequential when dependencies change. A model alias may resolve differently, a feature or data source may change, preprocessing and runtime configuration may evolve, and policy v8 may become v9. In agentic systems, retrieved context, memory, and tool contracts introduce further independent mutation. Deployment identity describes what was declared, but the historical record of an execution must describe what was actually used. Unless those binding relations are captured when they occur, the effective configuration of E891 cannot be recovered reliably from later system state. Figure 1. A deployment captures the initial configuration, while its execution-scoped dependency closure emerges as additional dependencies are bound during execution. (Diagram by the author.) Current state is a lossy projection Once runtime bindings are treated as part of the execution, the limitation of current-state records becomes clear. The execution-scoped dependency closure is the combination of the declared state and the runtime bindings that formed the effective configuration of a particular execution. As dependencies mutate, different historical executions can leave behind the same observable evidence. The current state may show which artifacts still exist without revealing which combination a particular execution used. Once the distinguishing bindings are gone, the history cannot be reconstructed from what remains. A system can retain its artifacts while losing reconstructability: An identifier proves neither that a dependency was used nor that the historical state it denotes remains available. Reconstructability therefore requires durable binding relations and continued access to the historical states they identify. Existence is not usage The existence of an artifact and its use in a particular execution are different facts. Concurrent versions, caches, retries, and asynchronous changes make it unreliable to infer whether an execution used an artifact merely because that artifact existed at the relevant time. W3C PROV already provides the relevant semantic distinction: An activity can use an entity. The missing primitive is therefore not a new provenance vocabulary, but a runtime requirement to record which dependency a particular execution used and to preserve an identity that can be resolved later. This requirement also exposes the limit of observability. OpenTelemetry can transport dependency identifiers through attributes, links, context, and baggage, but encoding an identifier does not preserve its historical meaning. A trace may retain execution topology while the identities that explain it disappear. Instrumentation can carry preservation metadata, but it cannot guarantee durable storage or continued resolution of the referenced states. Historical reconstructability therefore requires more than retained artifacts. For a defined class of executions and a specified retention period, the system must preserve two properties: binding integrity, which records which dependency versions and states an execution actually used, and resolution integrity, which keeps those recorded identities resolvable to the historical states they denote. Preservation requires failure semantics Once reconstructability is treated as a system property, the architecture must define what happens when the records required to reconstruct an execution fail to reach durable storage. Suppose policy v8 authorizes execution E891 to invoke tool contract h42, and the tool commits effect e3. If the usage record fails to reach durable storage, the action succeeds while the historical links among the effect, its policy, the tool contract, and the execution context are lost. The external effect remains, but the historical record no longer reliably explains which policy and tool contract authorized it. The external world and the historical record have diverged. Figure 2. Execution E891 branches into a committed external effect and a failed usage record. (Diagram by the author.) Figure 2 shows why this condition cannot be treated as ordinary telemetry loss. When an external effect has been committed but the corresponding usage record is missing, the architecture does not guarantee historical reconstructability. It provides only best-effort historical evidence. Reconstructability becomes an engineering guarantee only when the architecture defines when preservation records become durable and what the system does if they cannot be stored. A system may require a durable usage record before committing the effect, atomically persist the effect intent and preservation record through a transactional outbox, quarantine the execution for reconciliation, or trigger a compensating action. The implementation is application-specific, but preservation failure must have explicit semantics rather than disappear into telemetry. Existing technologies provide the building blocks for this layer. Provenance and lineage models represent relations, tracing propagates execution context, registries identify versions, and content-addressed stores preserve artifacts. None, however, creates a runtime obligation to preserve the execution-scoped dependency closure required for reconstruction. This layer is a preservation plane: the contract and machinery that keep an execution’s dependency closure identifiable as its surrounding systems change. It defines what to record, how records become durable, and how referenced states remain resolvable. Orrery applies this principle as a reference architecture in which materially relevant runtime bindings are captured when they occur and retained together with resolvable identities for the historical sta [truncated for AI cost control]