2026-06-30 04:00 UTCOriginal source2 min readUpdated: 2026-06-30 08:22 UTC

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

RoboGaze is a training-free, multi-agent VLM framework that provides structured, interpretable evaluation for generated robot-manipulation videos. It uses a three-stage pipeline and outputs localized glitch reports under a novel taxonomy, outperforming zero-shot baselines by large margins.

SourcearXiv RoboticsAuthor: Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem

[2606.28385] RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

[Submitted on 22 Jun 2026]

Title:RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

View a PDF of the paper titled RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis, by Minh-Loi Nguyen and 12 other authors

View PDF HTML (experimental)

Abstract:Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these videos is challenging: visually realistic outputs often violate physical laws, temporal consistency, or task logic, while conventional metrics and monolithic Vision-Language Model (VLM) judges fail to generalize or provide precise diagnostic value. We present RoboGaze, a training-free, multi-agent VLM framework that provides structured, interpretable evaluation for generated robot-manipulation videos. Given a task instruction and video, RoboGaze operates via a three-stage pipeline: task-scene grounding, dimension-specific specialist routing, and critic-based verification. It outputs temporally localized glitch reports categorized under a novel 6-dimension, 30-type robotics-specific taxonomy. To benchmark RoboGaze, we introduce a human-validated dataset of 382 clips spanning simulated and real-world multi-view manipulation. Evaluating eight open-source and proprietary VLM backbones, RoboGaze dramatically outperforms zero-shot baselines, improving description-F1 by up to +43 points and temporal alignment (F1 x IoU) by up to +37 points, closing approximately 85% of the gap to the human ceiling. Furthermore, its critic verifier mitigates the "cry-wolf" false-positive flaw of standard VLMs, lifting clean-clip accuracy from under 25% to over 80%. RoboGaze offers a scalable, highly interpretable diagnostic tool for the rigorous evaluation of robot world models.

Comments: First version 29 pages, 7 figures. Project webpage: this https URL

Subjects:

Robotics (cs.RO); Artificial Intelligence (cs.AI)

Cite as: arXiv:2606.28385 [cs.RO]

(or arXiv:2606.28385v1 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2606.28385

arXiv-issued DOI via DataCite

Submission history

From: Minh-Loi Nguyen [view email] [v1] Mon, 22 Jun 2026 06:45:09 UTC (37,091 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis, by Minh-Loi Nguyen and 12 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.RO

new | recent | 2026-06

Change to browse by:

cs cs.AI

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)