Skip to content
AI News HubLIVE
In-site rewrite2 min read

τ0-WM: The Largest Open-Source Embodied World Model Pre-trained on 17,800 Hours of Real Robot Data

Summary

The τ0-World Model (τ0-WM), a 5B-parameter open-source embodied world model, is pre-trained on nearly 30,000 hours of data, including 17,800 hours of real-world teleoperation data. It incorporates test-time computation to let robots simulate and evaluate multiple action sequences before execution, achieving state-of-the-art results on long-horizon manipulation tasks.

Source量子位Author: 衡宇
τ0-WM: The Largest Open-Source Embodied World Model Pre-trained on 17,800 Hours of Real Robot Data
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

A team led by Luo Jianlan, associate professor at Shanghai Innovation Institute and chief scientist at Agibot, has released τ0-World Model (τ0-WM), the largest open-source pre-trained embodied world model to date. With 5 billion parameters, τ0-WM was pre-trained on approximately 30,000 hours of data, including 17,800 hours of real-world teleoperation data—the first time such data has been used as the primary source for pre-training in embodied AI.

Unlike reactive end-to-end policies that directly output actions from visual input, τ0-WM enables robots to "think before they act" through test-time computation. The online inference pipeline consists of three stages: proposal, simulation, and evaluation. First, a video-action model (VAM) generates multiple candidate action chunks and corresponding future video latents. Second, an action-conditioned video simulator produces multi-view future frames for each candidate. Finally, a re-denoising consistency score (RCS) ranks actions, and if the top score is insufficient, a low-quality action rectification (LAR) mechanism triggers: the system selects the most promising future state and re-generates actions accordingly. This allows the robot to deliberate among alternatives before moving.

The training data is a carefully curated mixture. Real teleoperation data (17,800 hours) provides high-quality action supervision. UMI data (6,500 hours) adds behavioral diversity but with mismatched action spaces, so it contributes to both video and action training via modality-specific masks. Ego-centric human video (3,000 hours) lacks robot action labels and thus only trains the video branch, helping the model learn object dynamics and interaction patterns.

Experiments demonstrate the effectiveness of test-time computation. On two novel tasks (tissue-to-box and pen-to-box), the baseline policy without test-time computation achieved 43% average success rate. Adding RCS raised it to 50%, and further incorporating LAR brought it to 60%. In pen-to-box, the improvement was even more dramatic—from 30% to 50%. Comparisons with other guided methods (CFG 20%, ACG 38%) showed τ0-WM's advantage, as it evaluates future task progress rather than action consistency alone.

τ0-WM also reflects a paradigm shift in the data pyramid for embodied AI. Previously, real robot data was considered too expensive for pre-training and was reserved for fine-tuning. By combining large-scale teleoperation data with ego-centric video, the team has demonstrated that real robot data can serve as pre-training fuel, enabling a virtuous cycle of training, deployment, data collection, and retraining.

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • τ0-WM is the largest open-source pre-trained embodied world model with 5B parameters and ~30,000 hours of training data.
  • Real robot teleoperation data (17,800 hours) dominates the pre-training, a first in the field.
  • The model uses a propose-simulate-evaluate pipeline with test-time computation for deliberate decision-making.
  • It outperforms baselines on four long-horizon tasks, with test-time computation boosting success rates by 17 percentage points.

Highlights and analysis are generated automatically and may contain errors. Check the original source.