跳到主要內容
AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:Building an AI-powered contract intelligence platform with Amazon Quick and Amazon Bedrock AgentCore

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Manually extracting data from hundreds of vendor contracts doesn't scale, and RAG chat tools fall short on portfolio-wide questions. This post shares a contract intelligence platform on AWS that uses AI agents to extract and verify contract fields, then answers aggregate and single-contract questions through Amazon Quick analytics.

來源AWS Machine Learning Blog作者: Konala McGrath
待翻譯:Building an AI-powered contract intelligence platform with Amazon Quick and Amazon Bedrock AgentCore
回報錯誤

更正管道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

You’re a director of contracting, responsible for hundreds, maybe thousands, of vendor contracts. Each one is packed with critical data: contract values, expiration dates, signing status, and key contacts. With that information locked inside PDFs, you and your team spend hours manually extracting it, maintaining spreadsheets, and fielding the same recurring questions: “Which vendor are we spending the most with?” or “Which contracts are about to expire?” So you turn to AI chat tools and enterprise Q&A solutions. You upload a contract, ask questions in natural language, and for a single document, or even a handful, it genuinely works well. The real challenge shows up at scale, across hundreds of contracts. Most of these tools rely on a technique called Retrieval Augmented Generation (RAG). RAG breaks long text into chunks and indexes each one individually. During a search, RAG retrieves only the top-k chunks of text most relevant to a question, an approach more widely known as semantic search. That’s fine when the answer lives in one place. But when a question spans every contract, like total exposure or upcoming renewals, semantic search falls short. The answer requires aggregation across the full dataset, not only a handful of chunks. A better prompt won’t fix this. A different architecture will. In this post, we share the architecture behind a contract intelligence platform on AWS. It extracts contract data with AI agents, verifies accuracy with a multi-model approach, and delivers instant answers through embedded analytics and natural language querying. Everything runs from a single web application. The challenge: When AI chat tools hit their limit Say your portfolio holds 250 vendor contracts, and leadership keeps asking the same four questions: Total contract value across the portfolio. Which contracts have already expired. The most expensive agreement in the portfolio. The most recently signed deal. These are basic questions, but the answers are buried. Each of your 250 contracts runs 10–20 pages, so that’s up to 5,000 pages of unstructured data. To answer those questions, an analyst opens each PDF, searches for the fields that matter, and copies the values into a spreadsheet, then repeats the process 250 times. That’s more than a week of effort to clear the backlog, before a single new contract arrives. Every follow-up question from leadership means another round of filtering and analysis. Each round risks working from stale data or breaking hardcoded formulas on the new data. So you upload the portfolio to one of the leading AI chat tools and ask for the total value. The answer comes back fast, confident, and wrong. Where RAG fits, and where aggregation is needed That wrong answer wasn’t a fluke. It’s baked into how these tools are built. Internally they chunk each document into vectors and store them in a knowledge base. When you ask a question, they pull back the top-k chunks that best match it, then build an answer using only those chunks as context. For a targeted lookup, that’s exactly what you want. Ask “What are the payment terms in the AnyCompany contract?” and the right chunk surfaces with a clean answer. A portfolio question works against that mechanism instead of with it. When you ask, “What’s the total contract value across all 250 contracts?”, the system still returns only the top-k chunks and totals only those. It can’t sum, count, or compare across the portfolio because the portfolio never lands in front of the model. It isn’t a flaw in any one tool but in how RAG itself works. The key insight The answer is structured extraction, not better retrieval. To answer aggregation questions across hundreds of documents, you extract the key fields into a database, then query them with analytics tools built for exactly that. AI handles the unstructured-to-structured conversion at scale, and the database is used for math. You keep the original documents in a knowledge base for the single-contract questions it handles well. The result is a system that answers both kinds of questions: portfolio-wide aggregates and precise single-document lookups. Solution overview The platform is a React web application on AWS that turns a portfolio of contract PDFs into structured, queryable data. AI agents extract the key fields from each contract and independently verify them, with Amazon Textract settling any signature disagreements. Users then explore the results through embedded dashboards and a natural language chat agent, getting both portfolio-wide answers and single-contract details. The pipeline works as follows: You store contract PDFs in an Amazon Simple Storage Service (Amazon S3) bucket, which triggers an automated processing pipeline. An AI extraction agent (powered by Claude Sonnet series of models) reads the PDF and extracts eight key fields, each with a confidence score. A separate AI verification agent (powered by lighter Claude Haiku model) independently reads the same PDF and verifies the extraction. If the two agents disagree on signature detection, Amazon Textract provides a deterministic tiebreaker using computer vision. The system stores verified results in Amazon Aurora PostgreSQL. Users see real-time pipeline status over WebSocket and can immediately query data through embedded dashboards and a natural language chat agent. The entire pipeline can process a contract in seconds under typical conditions, and the serverless architecture is designed to help scale to handle many contracts in parallel. Here is the reference architecture for this solution. Figure 1: Contract intelligence pipeline, from ingestion to query Note: This post references the language models available when the solution was built. Model options change quickly, use the models available to you and re-test the solution before relying on the results. Inside the AI extraction pipeline The pipeline turns each uploaded contract into verified, structured data. It starts with two agents, one that extracts the fields and one that independently checks them, both running on fully managed infrastructure. The following section covers how they work together and the models behind them. Dual-model verification with Strands Agents on Amazon Bedrock AgentCore The extraction and verification agents are built with the Strands Agent SDK and deployed on AgentCore runtime, a capability of Amazon Bedrock AgentCore. The Strands Agent SDK is an open-source, model-driven framework for building autonomous agents. AgentCore handles serverless hosting, automatic scaling, and session isolation, so there’s no infrastructure to manage for the AI components. With Policy in Amazon Bedrock AgentCore, you can define and enforce security controls as a protective boundary around agent operations. This securely handles sensitive data such as pricing terms, financial commitments, and vendor relationships. Cedar-based policies exist outside agent code, and you can validate them with automated reasoning before enforcement. Agents and users only access the contracts they are authorized to see. We deliberately chose two different foundation models for extraction and verification: Claude Sonnet 4.6 for extraction: Advanced document understanding capabilities, reads the full PDF natively (no optical character recognition (OCR) preprocessing needed), and returns structured JSON with confidence scores for each field. Claude Haiku 4.5 for verification: Fast, cost-effective, and provides an independent perspective. Because it’s a different model with different training, it catches errors that re-running the same model would miss. For model availability by Region, refer to Supported models by AWS Region in Amazon Bedrock. This dual-model pattern is key to high accuracy on contract data, where mistakes are costly. A single model might hallucinate a value with high confidence. Two independent models disagreeing is a strong signal that human review is needed. Amazon Textract: The signature detection tiebreaker During testing, we found a revealing failure mode. The verification model would sometimes flag a contract as “signed” with 95–100% confidence when no signature existed, disagreeing with the extractor, which had correctly read it as unsigned. The model misread empty signature blocks, treating the presence of a signature field (“Signature: __________”) as proof of a signature. Rather than accepting this hallucination or writing elaborate prompts to work around it, we added Amazon Textract computer-vision-based signature detection as an architectural tiebreaker. Amazon Textract uses visual analysis, not language understanding, to detect actual handwritten or digital signatures on the page. It runs only when the two models disagree on the is_signed field, which keeps costs minimal while catching the false positives. This illustrates an important pattern: use large language models (LLMs) for what they excel at (document comprehension, field extraction, contextual understanding) and deterministic services for what they struggle with (visual element detection, precise counting, mathematical operations). Data-driven model selection The pipeline’s accuracy depends most on the model that handles extraction, so we tested several extractor and verifier combinations before settling on one. Using our 20-contract sample dataset, we hand-labeled all 8 fields (160 values total) to establish ground truth, then measured how closely each combination matched it. The evaluation surfaced two patterns. Extraction is the higher-impact step. A more capable extraction model held accuracy up even when paired with a lighter verifier, while a lighter extraction model brought accuracy down regardless of the verifier. More capability wasn’t always worth the cost. The strongest models didn’t meaningfully outperform a capable, lower-cost verifier on our data. Amazon Bedrock gives you a broad choice of models, and the right combination depends on your data. This was a small, directional test on 20 contracts, not an exhaustive evaluation, and results will vary with contract format and field complexity. We strongly recommend running your own evaluation against your benchmarks and success criteria, using the models available to you at build time. Querying at scale with Amazon Quick Extracting and verifying the data solves only half the problem. People still need to ask questions of it, from broad portfolio totals to the details of a single contract. Amazon Quick connects both the structured records and the original documents behind a single interface, so anyone can ask either kind of question. Within Amazon Quick, the Amazon Quick Sight capability powers the dashboards, while natural language querying and chat handle plain-language questions. Bridging structured and unstructured With contract data now extracted into a PostgreSQL database, we connected it to Amazon Quick so users can explore it through visual dashboards and natural language querying. The result is a single interface that answers both types of question: Aggregation queries (through structured data and Topics): “What’s our total contract value?” → “$50 million across 20 contracts.” Document-specific queries (through the knowledge base): “What are the payment terms in the AnyCompany contract?” → the exact clause, pulled from the source PDF. Embedded dashboards The Amazon Quick Sight dashboards are embedded directly in the React application using the Amazon Quick Sight Embedding SDK. Without leaving the application, users get the contract analytics they need without requiring exports or separate business intelligence (BI) tools. Rather than caching in SPICE (Super-fast, Parallel, In-memory Calculation Engine) and incurring extra costs, we query the data directly, so the dashboards are always live. The moment a contract finishes processing, it shows up. The dashboard includes several views: Overview: Key performance indi [truncated for AI cost control]

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級後設資料。
  • Manually extracting data from hundreds of vendor contracts doesn't scale, and RAG chat tools fall short on portfolio-wide questions. This post shares a contract intelligence platf…

技術影響

可能影響 Agent 架構、工具呼叫、工作流自動化和產品整合。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。