本文にスキップ
AI News HubLIVE
原典の内容 · 翻訳・分析待ち5 分で読了

翻訳待ち:Evaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore

記事の要約

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Multi-agent systems need deeper guarantees than fluent responses: they must select the right tools, respect constraints, and explain their decisions. Learn how to build a Strands-based multi-agent supply chain decisioning system and evaluate it with Amazon Bedrock AgentCore Evaluations using built-in, custom, and explainability evaluators.

ソースAWS Machine Learning Blog著者: Kanishk Mahajan
翻訳待ち:Evaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore
誤りを報告

訂正窓口はまだ利用できません。記事情報をコピーして保存できます。

訂正案内
本文へ

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

A critical challenge that emerges as multi-agent systems move from experimentation to production is making sure that these systems are consistently helpful, accurate, and explainable in real-world scenarios. Enterprises are increasingly adopting multi-agent systems to solve complex, real-world problems that require reasoning across data sources, tools, and business constraints. From supply chain planning to financial analysis and customer operations, these systems go beyond simple question answering. They coordinate multiple specialized agents to make decisions, execute workflows, and generate actionable recommendations. While large language models can generate fluent responses, enterprise applications require much deeper guarantees, where agents must follow instructions reliably, select the right tools, respect constraints, and provide clear reasoning behind their outputs. Amazon Bedrock AgentCore is a platform to build, connect, and optimize agents at scale, with any framework or model. Amazon Bedrock AgentCore Evaluations, a capability of Amazon Bedrock AgentCore, is designed to address this challenge as a fully managed capability for assessing agent performance across development and production, so teams can measure accuracy, task success, and behavior across multiple quality dimensions. Traditional evaluation approaches that focus only on model response quality are insufficient for agentic systems, where correctness depends on tool selection, workflow execution, and adherence to business constraints. In addition to evaluation, production deployment of agentic systems requires responsible AI controls. Amazon Bedrock Guardrails provides configurable safeguards such as content filtering, denied topic detection, and grounding validation that complement the evaluation framework. While evaluations assess agent quality after execution, Guardrails enforce safety constraints during execution. In this post, we focus on operationalizing this evaluation framework with Amazon Bedrock AgentCore Evaluations support for both built-in evaluators and custom evaluators. Built-in evaluators offer pre-defined assessments for common quality dimensions such as helpfulness, task success, and instruction following, so teams can quickly baseline agent performance without additional setup. However, enterprise use cases require deeper, domain-specific validation. Custom evaluators address this, so you can define business-aware checks. We also focus specifically on explainability as a first-class evaluation dimension. We demonstrate how built-in evaluators can assess general response clarity. We showcase how custom evaluators are used to verify that agents explicitly articulate decision rationale, reference supporting data or tool outputs, and explain tradeoffs such as cost versus service level. By combining these evaluators, we show how AgentCore Evaluations can move beyond surface-level response quality and provide structured, measurable insights into how and why agents arrive at their decisions. To make these concepts concrete, the following sections walk through a reference architecture and implementation that demonstrates how these components work together in practice. Solution overview For this post, we use a fictitious global retail company called AnyCompany Retail, a multinational retailer operating ecommerce channels, regional fulfillment centers, distribution centers, and thousands of physical stores. AnyCompany experiences frequent inventory imbalances: some regions face stockouts during promotions, while others carry excess inventory. Transportation teams must also balance delivery speed, carrier capacity, and cost. The company wants an agentic assistant that can help planners optimize inventory allocation, recommend distribution adjustments, analyze inventory health, and simulate routing or fulfillment scenarios. You will build and evaluate a multi-agent supply chain decisioning system using Strands Agents SDK, Amazon Bedrock AgentCore MCP Server and Amazon Bedrock AgentCore Evaluations. The solution uses Strands Agents with an orchestrator agent and four specialized sub-agents: an optimization agent, distribution agent, routing agent, and analytics agent. Each agent runs on Amazon Bedrock AgentCore runtime with Amazon Bedrock AgentCore memory and Amazon Bedrock AgentCore Observability enabled. The orchestrator agent receives the planner’s request and delegates work to specialized agents exposed as tools. The optimization agent calls MCP tools backed by mock Amazon API Gateway REST interfaces that return optimization decisions. The distribution agent calls recommendation APIs to suggest inventory rebalancing across fulfillment centers, stores, and digital channels. The routing agent calls logistics APIs to recommend carrier and route options and the analytics agent answers supply chain diagnostics questions. This solution uses foundation models on Amazon Bedrock for the agent loop. For model availability by Region, refer to Supported models by AWS Region in Amazon Bedrock. The solution uses built-in evaluators that assess general quality dimensions such as helpfulness and task completion. It also provides custom evaluators that assess supply-chain-specific behavior such as constraint satisfaction, route feasibility, SQL correctness, inventory grounding, and explanation quality. AnyCompany can evaluate both the language quality of the response and the business validity of the agent’s decision. The solution supports both on-demand and online modes with Amazon Bedrock AgentCore Evaluations. The on-demand mode is meant for development benchmarking, regression testing, and continuous integration and continuous delivery (CI/CD) gates. The online mode is for continuous production monitoring and alerts. Both modes help you close the loop and act on feedback from your users. The same custom evaluators (such as the constraint satisfaction, route feasibility, SQL correctness, and explainability evaluators from your supply chain solution) used for on-demand evaluations are repurposed with an OnlineEvaluationConfig object that references the Amazon Resource Names (ARNs) of the evaluators and specifies a sampling rate (for example, 1–10% of production traces) along with optional session filters. The service then automatically reads traces from AgentCore Observability, scores them and streams results to Amazon CloudWatch dashboards and alarms. In this post, you will use the on-demand mode to test the solution. The following architecture diagram illustrates the various components of our solution. Figure 1: Architecture of the multi-agent supply chain decisioning solution Evaluation framework In this post, you will use a three-layer evaluation approach for multi-agent systems that progressively builds enterprise trust. The approach follows a clear progression starting with built-in evaluators for general quality then adding custom evaluators for business accuracy and finally layering explainability evaluators for trust and auditability. The first layer uses built-in evaluators requiring no setup. We apply Helpfulness as a universal baseline plus a second agent-specific evaluator targeting each agent’s primary failure mode: Tool Selection Accuracy for the orchestrator, Response Relevance for optimization and distribution, Instruction Following for routing, and Faithfulness for analytics. The second layer adds custom evaluators encoding domain-specific business rules: constraint satisfaction for optimization, data grounding for distribution, route feasibility for routing, SQL correctness for analytics, and plan coherence for orchestration. These validate business validity: did the recommendation respect budget limits, use real inventory data, and produce operationally correct outputs? The following table maps the two built-in evaluators and custom evaluator selected for each agent that you will implement here. The second built-in evaluator targets each agent’s primary failure mode, while the custom evaluator encodes domain-specific business rules that validate operational correctness. Agent Built-in evaluators Custom evaluators Orchestrator agent Helpfulness; Tool Selection Accuracy Plan coherence evaluator: Did it combine sub-agent outputs into a valid, non-contradictory recommendation? Tool trajectory evaluator: Did it route to the correct sub-agent? Optimization agent Helpfulness; Response Relevance Constraint satisfaction evaluator: budget, inventory coverage (demand ≤ qty ≤ 2× demand), and warehouse capacity constraints. Key performance indicator (KPI) attainment evaluator: fill-rate/revenue improvement target met. Distribution agent Helpfulness; Response Relevance Recommendation groundedness evaluator: recommendation is grounded in current inventory/demand data. Risk impact evaluator: recommendation improves stockout/overstock risk. Routing agent Helpfulness; Instruction Following Route feasibility evaluator: route respects delivery window, cost, carrier capacity, and region constraints. Service level agreement (SLA) evaluator: expected delivery meets target service level. Analytics agent Helpfulness; Faithfulness SQL correctness evaluator: query matches user intent. Data-grounding evaluator: response is supported by Amazon Relational Database Service (Amazon RDS) query results. No unsupported claims evaluator. Explainability The third layer of our evaluation approach applies explainability evaluators as distinct, cross-cutting checks across the agents. These independently assess whether agents articulate decision rationale, cite supporting evidence from tool outputs, explain which constraints shaped the response, articulate trade-offs between competing objectives, clarify why specific sub-agents were invoked, and disclose assumptions when data is incomplete. By separating explainability into its own evaluation layer, we can independently measure transparency. A recommendation can be accurate but unexplainable (passing custom evaluators but failing explainability), giving teams actionable signals about whether agents need better reasoning articulation rather than better decision logic. The following table defines the six independent explainability evaluators that you implement here and that are applied across agents as a cross-cutting layer. These assess whether agents articulate reasoning, cite evidence, explain constraints and trade-offs, and disclose assumptions. They are measured separately from accuracy so teams can distinguish unexplainable-but-correct responses from well-explained-but-wrong ones. Evaluator Agents What it checks Decision rationale quality All Did the agent explain why it made the recommendation? Evidence attribution Analytics, Distribution, Routing Did it cite the data fields, API response, or SQL result used? Constraint reasoning Optimization, Routing Did it explain which constraints shaped the final response? Trade-off explanation Optimization, Distribution, Routing Did it explain the trade-offs among cost, service level, and inventory risk? Tool-use explainability Orchestrator Did it explain why each sub-agent or MCP tool was invoked? Assumption disclosure All agents Did it clearly state assumptions when data was incomplete? Prerequisites Before you deploy this solution, set up your development environment with the following tools. Install the AWS Command Line Interface (AWS CLI) Install the AWS Serverless Application Model (AWS SAM) CLI v1.100.0+ Install Docker v20.x+ Install Node.js v18.x+ Install Python v3.11+ Dependencies The Strands Agents implementation also needs to have the following dependencies that are packaged in the DockerFile: strands-agents # Strands Agents multi-agent framework strands-agents-tools # Strands agent tools and utilities requests # HTTP library for API calls bedrock-agentcore # Amazon Bedrock agent core functionality boto3 # AWS SDK for Python (Boto3) Deploy a [truncated for AI cost control]

要点と分析を開く

記事インテリジェンス

投資家上級

要点

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Multi-agent systems need deeper guarantees than fluent responses: they must select the right tools, respect constraints, and explain their decisions. Learn how to build a Strands-…

要点と分析は自動生成され、誤りを含む場合があります。原典をご確認ください。