AI News HubLIVE
站內改寫3 分鐘閱讀

待翻譯:How Amtrak is building the data backbone for its largest transformation in over 50 years

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:"Every new trainset is a data-generating asset. The intelligence platform that connects...

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

How Amtrak is building the data backbone for its largest transformation in over 50 years | Databricks Blog Skip to main content "Every new trainset is a data-generating asset. The intelligence platform that connects their signals is what makes the transformation compounding."—Prathibha Prabakaran, Senior Director, Enterprise Architecture, Amtrak Amtrak operates the largest passenger railroad network in the United States, with 21,000 track miles connecting communities across the country. The organization is now undertaking its biggest physical transformation in half a century, introducing two entirely new fleets while simultaneously rebuilding tunnels, bridges, and rail yards, some of which date back over 150 years. To match the scale of this physical transformation, Amtrak is building a digital intelligence platform that connects all of these new assets into a single governed layer. The team calls it Rail Intelligence, and it runs on Databricks. Siloed data and a once-in-a-generation opportunity Amtrak's data landscape has not kept pace with the speed of the transformation. Fleet telemetry lives in one system. The reservation platform, a 40-year-old mainframe called Arrow, lives in another. Capital project data sits in spreadsheets. Wayside detector readings are somewhere else entirely. Mechanical teams react to equipment failures after they happen. Analysts can't connect fleet health to crew scheduling to passenger demand. Meanwhile, new assets are coming online fast. The NextGen Acela, America's fastest train at 186 mph, is now in service with 28 trainsets on the Northeast Corridor. Eighty-three Siemens-built Airo trainsets are rolling out across 14 corridors. Each of these modern trainsets carries over 100 sensors generating thousands of data points per trip. Without a unified platform, that data would simply create new silos faster than the old ones can be retired. One platform for every signal from every train Amtrak selected Databricks as its strategic data platform. Not a point solution for a single use case, but the foundation for the organization's ability to scale as it goes through this massive transformation. Through Lakeflow Connect and real-time streaming, signals from across the enterprise now flow into a single governed environment: fleet IoT, wayside detectors, dispatch systems, geospatial feeds, the new Sqills S3 Passenger reservation platform, and operational technology systems. Raw events land in Delta Lake, get cleansed and conformed through a medallion architecture, and emerge as trusted data products governed by Unity Catalog with full lineage, domain access control, and data quality contracts. On top of this foundation, ML models run anomaly detection, a computer vision defect pipeline, and delay probability scoring, managed through MLflow and deployed via Model Serving. The principle is simple: one platform, every signal from every train, real time, governed, and trusted. Predictive maintenance, operational readiness, and smarter capital decisions The platform powers five intelligence products, each tied to a specific operational outcome. Fleet Health Intelligence streams continuous telemetry from Acela and Airo trainsets and surfaces predictive alerts for issues like door faults, bearing temperature deviations, and power car anomalies. Mechanical teams now get predictive signals rather than reacting to issues after the fact. Safety Intelligence automates ride quality monitoring and incident trend analysis, and proactively identifies food safety risks in café cars through refrigeration sensor data. Operational Readiness uses ML-powered scoring to unify fleet availability, crew scheduling, and maintenance windows into a single real-time operational picture. The goal: right train, right crew, right corridor, every departure, every day. Reservations Intelligence supports the migration from the legacy Arrow system to Sqills S3 Passenger, a cloud-native platform. Booking events, fare classes, and load-factor signals stream directly into the lakehouse via Lakeflow Connect. Capital Prioritization brings together wayside inspection data, ML anomaly scores, and fleet telemetry into unified condition scoring, so that Amtrak's $5.5 billion annual capital program is informed by live asset data rather than periodic manual assessments. From observable to compounding Amtrak's intelligence maturity is progressing in stages. The platform is live and governed today, with ML anomaly detection running across multiple fleets. The next stage is fully predictive, with delay probability models, revenue prediction, and capital scoring. The long-term vision includes agentic workflows and natural language queries through Genie that let operators ask questions of the data without writing code. Amtrak is also building a Databricks Apps experience layer to serve as a single entry point for developers, analysts, and executives to discover data products, view lineage from Unity Catalog, and interact with the platform through embedded Genie. The aim is to make data less abstract and more self-service across the organization. The compounding nature of the approach is what ties it all together. Each new trainset that enters service adds telemetry that improves every model. Each completed capital project adds condition data. Each passenger booking adds a demand signal. The platform grows more valuable with every asset that connects to it because the railroad that knows itself is the railroad that runs on time. Get the latest posts in your inbox Subscribe to our blog and get the latest posts delivered to your inbox. Sign up View all blogs