A team led by Luo Jianlan, associate professor at Shanghai Innovation Institute and chief scientist at Agibot, has released τ0-World Model (τ0-WM), the largest open-source pre-trained embodied world model to date. With 5 billion parameters, τ0-WM was pre-trained on approximately 30,000 hours of data, including 17,800 hours of real-world teleoperation data—the first time such data has been used as the primary source for pre-training in embodied AI.
Unlike reactive end-to-end policies that directly output actions from visual input, τ0-WM enables robots to "think before they act" through test-time computation. The online inference pipeline consists of three stages: proposal, simulation, and evaluation. First, a video-action model (VAM) generates multiple candidate action chunks and corresponding future video latents. Second, an action-conditioned video simulator produces multi-view future frames for each candidate. Finally, a re-denoising consistency score (RCS) ranks actions, and if the top score is insufficient, a low-quality action rectification (LAR) mechanism triggers: the system selects the most promising future state and re-generates actions accordingly. This allows the robot to deliberate among alternatives before moving.
The training data is a carefully curated mixture. Real teleoperation data (17,800 hours) provides high-quality action supervision. UMI data (6,500 hours) adds behavioral diversity but with mismatched action spaces, so it contributes to both video and action training via modality-specific masks. Ego-centric human video (3,000 hours) lacks robot action labels and thus only trains the video branch, helping the model learn object dynamics and interaction patterns.
Experiments demonstrate the effectiveness of test-time computation. On two novel tasks (tissue-to-box and pen-to-box), the baseline policy without test-time computation achieved 43% average success rate. Adding RCS raised it to 50%, and further incorporating LAR brought it to 60%. In pen-to-box, the improvement was even more dramatic—from 30% to 50%. Comparisons with other guided methods (CFG 20%, ACG 38%) showed τ0-WM's advantage, as it evaluates future task progress rather than action consistency alone.
τ0-WM also reflects a paradigm shift in the data pyramid for embodied AI. Previously, real robot data was considered too expensive for pre-training and was reserved for fine-tuning. By combining large-scale teleoperation data with ego-centric video, the team has demonstrated that real robot data can serve as pre-training fuel, enabling a virtuous cycle of training, deployment, data collection, and retraining.