AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
When a user asks the support assistant, a Retrieval Augmented Generation (RAG) application built with LangChain to compare two products across three dimensions, they’re effectively posing six questions simultaneously. Similarity search uses a single query vector to encapsulate all the intents. The retriever then generates the best approximation of the average of those intents. The resulting answer comes back concise. The search executes without errors. The relevance scores look reasonable. Yet the retrieved chunks, while topically relevant, only cover a fraction of what the question actually asked. In this post, we showcase a RAG application on Amazon Bedrock Managed Knowledge Base with LangChain. We run the same multi-part question through standard and agentic retrieval, and read the trace events to see the plan the model produced. We also cover what the two retrieval paths cost and when the cheaper one is the right choice. Agentic retrieval is available on Amazon Bedrock Managed Knowledge Base. Instead of one search, Amazon Bedrock Managed Knowledge Base plans the retrieval. It breaks the question into sub-queries, runs them, judges whether it has enough evidence, and searches again if it doesn’t. The langchain-aws package exposes both agentic and standard retrieval, so you can use either from a LangChain application. Solution overview Amazon Bedrock Managed Knowledge Base, the fully managed RAG capability in Amazon Bedrock, removes the self-managed vector store, embeddings, and re-ranking models from the RAG architecture. You configure a data source, and Amazon Bedrock Managed Knowledge Bases handles chunking, embedding, storage, and retrieval. This walkthrough uses Amazon Simple Storage Service (Amazon S3). Amazon Bedrock Managed Knowledge Bases provides two APIs. We briefly discuss those differences in this post. The Retrieve API runs one hybrid search and returns scored chunks. The AgenticRetrieveStream API runs a planning loop and streams the steps back to you as trace events. In the langchain-aws package, the first is a standard LangChain retriever you can drop into a chain. The second is a function retrieval directly from a knowledge base. The following diagram shows the solution architecture. The application queries Amazon Bedrock Knowledge Bases using either the Retrieve API (standard, single-shot) or the AgenticRetrieveStream API (multi-step planning loop). Both paths return document chunks from the knowledge base, which the application then uses to generate a grounded response. Figure 1: Solution architecture for querying Amazon Bedrock Knowledge Bases with the Retrieve and AgenticRetrieveStream APIs Implementation walkthrough The following sections walk you through creating a knowledge base, querying it with both retrieval methods, and reading the trace events the agentic planner produces. Prerequisites To follow along you need: An AWS account with access to Amazon Bedrock in a Region where Amazon Bedrock Managed Knowledge Bases and agentic retrieval are available. This walkthrough uses the US East (N. Virginia) Region (us-east-1), and the code assumes it throughout. Check the AWS documentation for other Regional availability and support. Two AWS Identity and Access Management (IAM) identities, described in the next section: a service role the knowledge base assumes, and permissions on the identity you call the APIs from. Python 3.12 or later. An S3 bucket holding the sample documents. The corpus needs several documents that cover overlapping topics so that a comparative question has somewhere to go. A single flat document cannot demonstrate query planning. Install the packages. The Boto3 version matters: agentic_retrieve_stream did not exist before 1.43.32. langchain-aws>=1.6.3 langchain>=1.0 boto3>=1.43.32 Permissions Two identities are involved and separating them is worth doing deliberately. The knowledge base assumes a service role to read your documents and call the embedding model. Your application uses an AWS Security Token Service (AWS STS) caller identity to query. Neither needs the other’s permissions. Amazon Bedrock creates the service role for you if you let it. To supply your own, give it a trust policy that lets Amazon Bedrock assume it. Scope it with aws:SourceAccount and aws:SourceArn so that another account can’t use it as a confused deputy: { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Principal": {"Service": "bedrock.amazonaws.com"}, "Action": "sts:AssumeRole", "Condition": { "StringEquals": {"aws:SourceAccount": "111122223333"}, "ArnLike": { "aws:SourceArn": "arn:aws:bedrock:us-east-1:111122223333:knowledge-base/*" } } }] } The service role also needs s3:ListBucket on your bucket and s3:GetObject on its contents, both conditioned on aws:ResourceAccount. Scope the knowledge-base/* wildcard character down to specific knowledge base IDs after you have created them. The AWS STS caller identity needs a different set. bedrock:AgenticRetrieveStream and bedrock:InvokeModelWithResponseStream can’t be scoped to a knowledge base Amazon Resource Name (ARN). bedrock:Retrieve and bedrock:GetDocumentContent can: { "Version": "2012-10-17", "Statement": [ { "Sid": "AgenticRetrievalAndPlannerModel", "Effect": "Allow", "Action": [ "bedrock:AgenticRetrieveStream", "bedrock:InvokeModelWithResponseStream" ], "Resource": "*" }, { "Sid": "RetrieveAndFullDocumentExpansion", "Effect": "Allow", "Action": ["bedrock:Retrieve", "bedrock:GetDocumentContent"], "Resource": "arn:aws:bedrock::111122223333:knowledge-base/" }, { "Sid": "GenerateAnswersInTheChains", "Effect": "Allow", "Action": ["bedrock:InvokeModel", "bedrock:Converse", "bedrock:ConverseStream"], "Resource": "*" } ] } bedrock:GetDocumentContent is often overlooked. Agentic retrieval calls it when a FullDocumentExpansion step decides a passage lacks the context to answer. A policy with only bedrock:Retrieve works until the planner reaches for a whole document and then fails partway through a query. To create and manage the knowledge base itself, the calling role additionally needs bedrock:CreateKnowledgeBase on *, and the GetKnowledgeBase, UpdateKnowledgeBase, DeleteKnowledgeBase, StartIngestionJob, GetIngestionJob, and ListIngestionJobs actions on knowledge-base/*. If you’re using guardrails, add bedrock:GetGuardrail and bedrock:ApplyGuardrail. Running this walkthrough might incur costs for document storage and ingestion in the knowledge base, retrieval calls, and foundation model (FM) inference. For more information about pricing, see the Knowledge Bases section of Amazon Bedrock pricing. Delete the resources when you complete this experiment. Creating and populating the knowledge base Create the knowledge base with a managedKnowledgeBaseConfiguration. Setting embeddingModelType to MANAGED uses the service-managed embedding model. import boto3 import os REGION = os.environ["AWS_REGION"] bedrock_agent = boto3.client("bedrock-agent", region_name=REGION) response = bedrock_agent.create_knowledge_base( name=KB_NAME, roleArn=KB_ROLE_ARN, knowledgeBaseConfiguration={ "type": "MANAGED", "managedKnowledgeBaseConfiguration": { "embeddingModelType": "MANAGED", }, }, ) KB_ID = response["knowledgeBase"]["knowledgeBaseId"] There’s no storageConfiguration in that request. For a self-managed knowledge base you would pass one describing your vector store. Amazon Bedrock Managed Knowledge Base does not take one, which is the clearest signal in the API that Amazon Bedrock owns the storage layer. Attach the S3 bucket as a data source, then start an ingestion job. Ingestion is asynchronous, so poll until the job reaches a terminal state rather than sleeping for a fixed interval and hoping. import time SUCCESS_STATES = frozenset({"COMPLETE"}) FAILURE_STATES = frozenset({"FAILED", "STOPPED"}) def wait_for_ingestion(kb_id, ds_id, job_id, timeout_s=1800): """Poll an ingestion job until it reaches a terminal state.""" deadline = time.time() + timeout_s while time.time() str: result = agentic_retrieve( knowledge_base_id=KB_ID, query=question, region_name=REGION, number_of_results=10, ) return "\n\n".join( item.get("content", {}).get("text", "") for item in result.get("results", []) ) agentic_chain = ( {"context": RunnableLambda(agentic_context), "question": RunnablePassthrough()} | prompt | llm | StrOutputParser() ) Note that generate_response is off here. The service can generate the answer itself, but inside a chain you usually want your own prompt and model, so you take the chunks and generate downstream. Use the service generation when you want one call and less code, and the wrapped version when the prompt is yours to control. Choosing between standard and agentic retrieval Use Retrieve for short, well-scoped questions. It is cheaper, faster, works against self-managed knowledge bases, and returns scores in the results. Most production traffic looks like this. Use AgenticRetrieveStream when questions are multi-part, comparative, or exploratory, or when the evidence spans more than one knowledge base. It registers up to five knowledge bases in one request and routes sub-queries using a natural-language description you attach to each. The other API cannot do this at all. It costs more per call, makes several model invocations, and has the higher latency of the two. Routing on query shape rather than picking one for everything is the pattern we recommend. A classifier or a heuristic on the question can send most traffic down the cheap path and reserve the planner for questions that need it. Clean up resources Delete the knowledge base, its data source, the S3 objects and bucket, and the IAM role that you created. A knowledge base with documents in it continues to incur storage charges. bedrock_agent.delete_data_source(knowledgeBaseId=KB_ID, dataSourceId=DS_ID) bedrock_agent.delete_knowledge_base(knowledgeBaseId=KB_ID) The repository includes a cleanup script that also empties the bucket and removes the role. Conclusion We showed how to build a RAG application on Amazon Bedrock Knowledge Bases with LangChain, and how agentic retrieval handles multi-part questions that single-shot retrieval answers poorly. We also showed the friction in the current integration. Agentic retrieval is a function rather than a LangChain retriever, so it needs a RunnableLambda to sit in a chain. The trace events that show the query plan require a direct boto3 call. Agentic retrieval trades higher per-call cost for improved recall on multi-hop questions, using a built-in model for query planning. The next useful step is measuring your own query mix before you route everything through a planner. To get started, see the Amazon Bedrock Knowledge Bases documentation and the accompanying sample code. For help applying this to your own workload, contact your AWS account team. About the authors