AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。
Checking tens of thousands of apartment leases against constantly changing state landlord-tenant laws, and proving you actually checked all of them, has been beyond the reach of most compliance teams. But with generative AI in Amazon Quick, paired with the right backend, it’s now possible. In this post, we introduce a design pattern called Adjudicated Query. Business users can use it to ask compliance questions in a chat interface (Amazon Quick), while the actual pass/fail decisions stay in a deterministic (non-AI) rules engine. We walk through an AWS reference architecture that implements the pattern, and deploy a working sample you can run end to end. The pattern applies to other high-stakes compliance domains as well (sanctions screening, insurance claims adjudication, export control), but lease compliance serves as our concrete example. The compliance challenge at scale A portfolio operator holds 50,000 leases across multiple states. Each state publishes landlord-tenant statutes (late-fee caps, notice periods, security-deposit limits) that change on the legislature’s schedule, not the operator’s. When a regulation changes, the team responsible for compliance must determine which leases are now out of line. At small volume a paralegal reads the leases. The answer is trustworthy because a human stands behind it. Past some threshold, that stops being possible. The work moves to software, and a new problem appears: the answer is now a number on a screen that nobody can independently verify. Two properties follow from that reality: Provable completeness: A claim like “we checked all 22,910 Texas leases” must be true and demonstrable. A record never assessed must be reported as unevaluated rather than silently omitted. Defensibility: A finding may be challenged months later in litigation, an audit, or a regulatory examination. Defending it means knowing which version of which rule was applied, to which clause text, by what method, on what date, and by whom. These two properties are what distinguish this problem from enterprise search. Retrieval Augmented Generation (RAG) addresses the accessibility gap but cannot satisfy either property. Similarity search has no threshold that means all of them. A ranked sample never knows what it excluded. Text-to-SQL narrows this gap, but carries a category-level risk: a hallucinated predicate can silently reduce the population, and the resulting number looks exact even when the scope is wrong. How the Adjudicated Query pattern solves it The Adjudicated Query pattern is a bounded conversational layer over a deterministic rules engine. The model does exactly two things: translate a natural-language question into a call on a fixed set of typed operations, and narrate the result that comes back. It never writes a query, never fixes the population, and never performs a determination. Behind the boundary sits a rules engine. Rules are versioned data, not code. The engine knows generic comparison operators (gte, lte, equals, exists) and contains no branch naming a jurisdiction or topic. A law change is a rulebook row edit, not a code deployment. Every compliance sweep produces a completeness receipt: an asserted invariant where compliant + in-breach + ambiguous + unreadable must equal scanned. This is computed from counts and asserted before anything persists. A run that can’t account for its population never finishes. There’s no path by which a record is silently skipped. The conversational surface carries counts, the receipt, and a labeled sample. The full result set (potentially tens of thousands of rows) lives on a dashboard surface reading the same data store, drillable per record. This separation means the model never summarizes away the guarantee. Why not RAG or text-to-SQL? Approach Population Completeness Defensibility Semantic retrieval (RAG) A ranked sample Structurally impossible Partial Generated queries (text-to-SQL) Claimed but unprovable Silent narrowing risk If modeled Rules engine + BI (no chat) Exact and proven Yes Yes Adjudicated Query Exact and proven Yes Yes The Adjudicated Query pattern adds natural-language access to the rules engine plus business intelligence (BI) approach without sacrificing the guarantee. It’s the right choice when accountable users need conversational access, and a missed record is a liability rather than a mild inconvenience. Reference architecture The following diagram shows how the components fit together end to end. A compliance officer interacts with two surfaces in Amazon Quick: a chat agent for asking questions and an Amazon Quick Sight dashboard for browsing the full result set. The chat agent first fetches an OAuth token from Amazon Cognito. It then sends Model Context Protocol (MCP) requests over an Amazon API Gateway HTTP API, which validates the token before forwarding to an AWS Lambda function. The Lambda function hosts the MCP server and the rules engine, reads and writes to Amazon Aurora Serverless v2 through the RDS Data API, and calls Amazon Bedrock only for the exploratory clause-search path. The Amazon Quick Sight dashboard reads the same Aurora store directly through a virtual private cloud (VPC) connection. Both surfaces therefore read from one store, which is what makes the completeness receipt a single source of truth. Figure 1: Reference architecture for the Adjudicated Query pattern Both surfaces read the same store. The chat agent carries the completeness receipt and a link to the dashboard. The dashboard carries the volume, because 10,800 rows don’t render in a chat message. Amazon Bedrock is called from AWS Lambda only by the exploratory operation. No model is involved in compliance sweeps, and Amazon Aurora Serverless v2 doesn’t call a model. The bounded operation surface The MCP server exposes exactly six tools, each with a distinct semantic: Tool What it does Result means sweep_compliance Exhaustive population sweep against rules in force on a stated date Official. Every lease accounted for in a computed receipt. Writes findings simulate_rule_change One rule tested at a proposed value against the approved baseline Exploratory. Directional counts only. Records nothing explore_clauses Top K by semantic similarity within a filtered population Interpretive. A ranked sample. Cannot answer “how many” get_finding One finding’s complete evidence chain Drill-down into a single determination list_rules The rulebook in force on a date, with versions, citations, approvers Reference lookup check_connection Liveness check, touches no data Transport health This bounded surface removes the path to the silent-narrowing risk of generated queries. Because the model can only select from a fixed set of operations whose population logic was written, reviewed, and tested by people, it has no way to compose a wrong population. Key architecture components With the Amazon Quick conversational interface and agent orchestration layer, you can ask natural-language compliance questions that the Amazon Quick chat agent translates into calls on the bounded MCP operation surface. Amazon Quick authenticates to the MCP server by using OAuth 2LO through Amazon Cognito and handles tool discovery and response narration. The deterministic engine handles the compliance logic. Amazon Aurora Serverless v2 (Postgres + pgvector) stores the rulebook, lease records, extraction status, determinations, and runs in a single relational store. Putting everything in one database makes the completeness receipt a SQL count, a cost-effective way to make the central guarantee inspectable. AWS Lambda hosts the MCP server (JSON-RPC 2.0 over Streamable HTTP, using Server-Sent Events framing for responses, which the Amazon Quick client requires) and the rule engine. Bounded operations translate to set-based SQL by using rule values bound as parameters. No natural language reaches the query layer. Amazon API Gateway HTTP API provides the front door with a JSON Web Token (JWT) authorizer backed by Amazon Cognito. No unauthenticated route exists. Amazon Cognito issues OAuth tokens through a two-legged (client credentials) flow. The client secret is read from Amazon Cognito at registration time and not written to disk. Amazon Bedrock powers the exploratory path only, using Amazon Titan Text Embeddings V2 (amazon.titan-embed-text-v2:0) for semantic similarity ranking and Anthropic Claude Sonnet 5, invoked through a cross-region inference profile, for qualitative clause assessment. It isn’t consulted for an official compliance determination. Amazon Bedrock model availability, including Amazon Titan Text Embeddings V2 and Anthropic Claude Sonnet 5, varies by AWS Region, so confirm the models are available in your Region before deploying. Amazon Quick Sight connects to Aurora through a VPC connection and renders the full findings table, filterable by sweep and severity band, with every column needed to defend a determination already on the row. Design rules that are non-negotiable Rules are data. A law change is a rulebook row. The engine contains no jurisdiction-specific branch. No natural language reaches SQL. Operators select fixed SQL templates. Rule values bind as parameters. Deterministic before AI. A numeric comparison answers the sweep. No model is consulted. Exact filtering for completeness, vectors only for ranking. Similarity never decides membership. Receipts are computed, not written by hand. The invariant is asserted before a sweep commits. Unreadable documents are named, not dropped. Every lease lands in exactly one bucket. Findings are append-only. No UPDATE or DELETE against findings exists in the code base. Treating the summarizing model as an untrusted renderer One design element deserves its own section because it will look unfamiliar: engineering safeguards to survive paraphrase by the chat model. Excluding the model from the decision path but putting one back in the delivery path reintroduces risk at the end of the chain. In practice, we observed: A model stripped a caveat prefix. An ILLUSTRATIVE citation tag was removed during paraphrase, and the model presented an invented citation as statute. A model extrapolated from a sample. Given 20 preview rows, the model inferred a population-wide range that did not exist in the data. Three techniques address this: Phrase caveats as un-strippable bracketed suffixes repeated at several payload levels, not leading labels that read as removable metadata. Supply the data that makes the honest answer the easy one. Compute real aggregates over every record and hand them to the model. A model that has the real number doesn’t need to guess from a sample. Repeat mode labels at multiple structural levels (field, string, summary) so that at least one survives paraphrase. The principle: a safeguard in the payload is only as strong as its survival through paraphrase. Deploy and run the sample The complete reference implementation is available on GitHub. It ships with synthetic data (no real customer lease data), deterministic corpus generation, and acceptance tests against the deployed stack. Prerequisites Before deploying, verify that you have: An AWS account with Amazon Bedrock model access enabled for Amazon Titan Text Embeddings V2 and Anthropic Claude Sonnet 5 in the US East (N. Virginia) Region (us-east-1). AWS Command Line Interface (AWS CLI) v2 with credentials configured (aws sts get-caller-identity should succeed). Python 3.12. Node.js 24 for the AWS Cloud Development Kit (AWS CDK) CLI. mise is optional and only used to provision Node 24. You can install Node 24 by your choice of method (Node 18 has reached end of life for CDK). Step 1: Clone the repository Clone the sample repository and set up the Python environment: git clone https://github.com/aws-samples/sample-quick-adjudicated-query.git cd sample-quick-adjudicated-query python3 -m venv .venv && .venv/bin/pip instal [truncated for AI cost control]