AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。
Trane Technologies manages millions of connected heating, ventilation, and air conditioning (HVAC) assets worldwide, but getting a single operational answer could mean cross-referencing multiple dashboards and drilling through menus for 20 minutes or more. For organizations operating at this scale, that kind of friction slows operations, defers corrective action, and creates material business impact across the enterprise. In 3–4 weeks, Trane’s engineering team built an AI-powered agentic solution on Amazon Bedrock AgentCore that reduced a 20-minute multi-screen diagnostic workflow to a 20-second natural language interaction. This is based on Trane’s internal benchmarking with technicians over several weeks. This represents a 60x improvement in time-to-insight, helping shift operations from reactive response to more proactive, data-driven optimization. In this post, we describe the architectural approach and key design decisions behind the solution: Separating agent logic from tool execution. Integrating real-time telemetry through a centralized tool gateway. Tailoring responses to different personas. Trane Technologies and the building intelligence challenge Trane Technologies is a global climate innovator with over $21 billion in annual revenue and operations in more than 100 countries. Through its strategic brand Trane, the company manages millions of connected HVAC assets, spanning data centers, hospitals, manufacturing facilities, and commercial real estate portfolios. At the heart of this vast landscape is Trane Cloud, a digital hub that aggregates real-time performance data from millions of HVAC systems. Trane Cloud transforms raw equipment telemetry into actionable intelligence for predictive maintenance, energy optimization, and operational excellence. While dashboard-based building management systems provide a foundation for monitoring and control, extracting cross-system insights can still require users to navigate multiple screens, layered menus, and disconnected dashboards. By combining natural language processing with deep integration into Trane Cloud, users can access that operational context through a single conversational interface. The result is faster answers to complex building management questions and a more proactive, informed approach to facility operations. Business challenge: Why building operators need an AI agent Building operators, field technicians, and service managers have abundant data at their fingertips. Equipment telemetry, performance analytics, fault alerts, energy consumption patterns, and optimization opportunities flood in from disparate systems, yet extracting actionable insights remains difficult. The fundamental problem is that different stakeholders need radically different views of the same data. Field technicians require diagnostic precision. They need refrigerant pressures, fault codes, and system-level troubleshooting workflows. Account managers need strategic intelligence. They need uptime metrics, cost savings opportunities, and portfolio performance trends. Building owners demand executive clarity. They need efficiency scores, sustainability metrics, and simplified operational summaries. Existing tools present a single interface across all roles, requiring each user to navigate features outside their workflow. Existing building management applications rely on screen-by-screen navigation that makes cross-equipment comparison more challenging and demands that users memorize menu hierarchies and technical terminology. Even basic portfolio-level questions can require time-intensive manual workflows across multiple screens. Together, Trane’s agentic AI solution and Trane Cloud deliver four capabilities: Role-based access control that tailors responses to each user’s permissions and needs. Real-time HVAC analytics providing instant access to current and historical performance data. Intelligent search across Trane’s knowledge base of technical documentation and best practices. Extensible architecture that evolves with advancing AI capabilities and expands to additional building systems. These capabilities create a more scalable way to access building intelligence across roles, workflows, and operational environments. Solution overview: Architecture The challenge of scaling intelligent building operations lies in turning the massive volume of data generated by millions of connected assets into actionable insight. Although Trane Cloud ingests real-time telemetry at scale, answering operational questions has traditionally required users to navigate disconnected dashboards and manually connect information across systems. To solve this, the team built a conversational agent on Amazon Bedrock AgentCore and the Strands framework, deployed through AWS Cloud Development Kit (AWS CDK) infrastructure as code. The team chose Strands for the agent framework layer because it provides the developer SDK and orchestration logic for building agent behavior, while AgentCore handles the managed runtime, memory, tool gateway, and production infrastructure underneath. To avoid the limitations of a monolithic design, the solution uses a multi-agent architecture where each specialized assistant is governed by its own system prompt, keeping it tightly focused on a single capability domain: Resources Assistant – Retrieves and summarizes reference material. For example, a user might ask: “Who do I contact, what manual should I follow, or what documentation can I share with the customer?” Knowledge Assistant – Synthesizes technical answers about how equipment works, its system parameters, or whether Trane Cloud’s infrastructure is SOC 2 attested. Analytics Insights Assistant – Interrogates live telemetry to surface efficiency opportunities, flag items needing inspection, and trace fault root causes. Expert Advisor – Helps users decide which product solution fits a scenario, how to maximize customer value, or how to assemble a customized demo. Navigation Assistant – Returns the exact links and tools a user needs, from the tech support escalation form to the replacement-parts order page. This architecture is designed for extensibility. Teams can connect additional agents or tools, such as work order management systems and enterprise customer relationship management (CRM) systems, through AgentCore Gateway, a capability of Amazon Bedrock AgentCore, and open standards like the Model Context Protocol (MCP). Figure 1: High-level architecture of Trane’s conversational agent on Amazon Bedrock AgentCore Microservices architecture: System design and implementation challenges Supporting these distinct user needs at enterprise scale requires an architecture that can evolve independently across capabilities. A monolithic agent would force every change (new tools, updated prompts, additional data sources) through a single deployment pipeline, creating bottlenecks as the system grows. Reducing a 20-minute manual diagnosis to a 20-second conversation surfaced four architectural challenges. The first challenge was integration. The solution had to combine real-time telemetry with intelligent search across an extensive knowledge base while supporting connections to external systems like CRMs. The second was separation. Agent logic had to be untangled from backend tool execution so the two could deploy independently, with clear ownership boundaries. Third was context. The system needed to maintain conversational state across troubleshooting sessions without standing up complex custom vector database infrastructure. Fourth was observability. When an agent orchestrates multiple tools across a multi-step reasoning chain, failures become difficult to localize. A wrong answer could stem from a missing API credential, a malformed tool response, or a model hallucination. Without end-to-end tracing, the team had no way to distinguish between them at production scale. How Amazon Bedrock AgentCore addresses the challenges Amazon Bedrock AgentCore is an agentic platform to build, connect, and optimize agents at scale, with any framework or model. The Trane team used four AgentCore capabilities to address the preceding challenges. Before selecting AgentCore, the team evaluated hosting the agent on Amazon Elastic Container Service (Amazon ECS) and AWS Lambda. That approach would have required building session isolation, auto scaling logic, and per-session billing on top of the compute layer. Four differentiators drove the decision. First, the managed agent runtime alleviates infrastructure operations. There are no clusters to provision or scale, and no idle capacity to pay for between user sessions. Second, built-in session memory removes the need to stand up and maintain external vector databases or build custom context-window management code. Third, native tool orchestration through AgentCore Gateway turns existing internal APIs into agent-compatible tools without writing custom integration logic for each one. Fourth, AgentCore’s framework-agnostic design meant the team could use the Strands SDK without being locked into a proprietary orchestration layer, preserving flexibility as requirements change. AgentCore runtime: Trane uses AgentCore runtime, a capability of Amazon Bedrock AgentCore, to isolate each user session in a dedicated microVM with its own CPU, memory, and filesystem. AgentCore runtime terminates and sanitizes each microVM on session completion. The team chose it because the microVM model separates the user-facing agent from the backend MCP Server, letting the two deploy independently with clear ownership boundaries. AgentCore runtime physically isolates a field technician’s session from a building owner’s, reinforcing role-based access without custom infrastructure. Trane pays only for active compute during a session. The long pauses between tool calls (typical of agentic workflows) don’t accumulate cost. AgentCore Gateway: Trane uses AgentCore Gateway to expose Trane Cloud’s internal APIs as MCP-compatible tools that the team organized by capability domain and integrated with Amazon OpenSearch Service. The team chose it because connecting the suite of assistants to real-time analytics, equipment telemetry, and issue diagnostics required a single access layer that handles authentication and schema translation. Future integrations (CRMs, work order systems) connect through the same Gateway without additional plumbing. AgentCore memory: Trane uses AgentCore memory, a capability of Amazon Bedrock AgentCore, to maintain conversational state across troubleshooting sessions so users can ask natural follow-ups (“now compare that to last month”) without re-specifying context. The team chose it because the alternative was standing up a separate vector database and writing custom context-window management code. AgentCore memory provides short-term session memory out of the box, with a 90-day expiry lifecycle that balances contextual awareness with storage efficiency, helping Trane address their data retention policies. AgentCore Observability: Trane uses AgentCore Observability, a capability of Amazon Bedrock AgentCore, to trace tool calls an agent makes and isolate failures across multi-step reasoning chains. The team chose it after encountering consistent tool failures during development that could not be identified without end-to-end visibility. Through the agent traces in Amazon CloudWatch, engineers confirmed the identical failure pattern across multiple invocations and traced it to a missing secret in AWS Secrets Manager. After adding the secret, the failures resolved. The real-time data flow works as follows: The user’s query, carrying a JSON Web Token (JWT), hits AgentCore runtime. The Runtime’s inbound authorizer (AgentCore Identity, a capability of Amazon Bedrock AgentCore) validates the token against the OpenID Connect (OIDC) discovery endpoint. AgentCore memory then injects previous conversation [truncated for AI cost control]