Skip to content
AI News HubLIVE
Original source2 min read

R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models

Summary

Evaluating robot manipulation policies is increasingly important as generalist models and VLA models are deployed on physical robots. This paper proposes R2S-Eval, which combines real-to-sim calibration with vision-language model (VLM) preference evaluation. Rollout videos are generated in a simulator calibrated to real-world settings to reduce repeated hardware trials, and a VLM assesses execution quality in pairwise comparisons that are aggregated into policy rankings. Experiments show more reliable and stable conclusions than conventional success-rate evaluation, agreement with human preferences, reduced hardware effort, and the ability to reveal behavior-quality differences invisible to binary success labels.

SourcearXiv RoboticsAuthor: Yidi Wang, Feixiang Ruan, Ruoqu Chen, Jie Yin, Yang Yu, Mengdi Xu, Kaifeng Zhang
R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

[Submitted on 3 Sep 2026]

Title:R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models

View a PDF of the paper titled R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models, by Yidi Wang and 6 other authors

View PDF HTML (experimental)

Abstract:Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly vision-language-action (VLA) models, are deployed on physical robots. However, conventional real-world evaluation remains labor-intensive, unstable, and insufficiently informative. It requires repeated hardware trials, manual scene resets, and continuous operator monitoring, may produce different policy rankings across repeated evaluations, and primarily relies on success-rate metrics that provide limited information about execution quality. In contrast, humans assess robot performance by observing and comparing complete behaviors rather than relying solely on binary success outcomes. To this end, we propose R2S-Eval, an evaluation pipeline that combines real-to-sim calibration with vision-language model (VLM) preference evaluation. The real-to-sim component efficiently generates rollout videos in a simulator calibrated to the real-world evaluation setting, thereby reducing the need for repeated hardware trials. The VLM evaluator assesses the execution quality of rollout videos and produces pairwise preferences, which are subsequently aggregated into policy rankings. We further introduce a protocol to assess whether the proposed evaluation pipeline yields validated policy conclusions while mitigating the key challenges of conventional real-world evaluation. Experiments in both simulation and real-world settings demonstrate that R2S-Eval produces reliable and stable policy conclusions, achieves agreement with human preferences, substantially reduces repeated hardware-operation effort, and reveals behavior-quality differences that are not captured by binary success labels. In general, R2S-Eval advances robot evaluation from manual success counting toward automated, statistically stable, and quality-aware evaluation of robot behavior. Project page: this https URL.

Subjects:

Robotics (cs.RO)

Cite as: arXiv:2609.03276 [cs.RO]

(or arXiv:2609.03276v1 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2609.03276

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yidi Wang [view email] [v1] Thu, 3 Sep 2026 02:08:16 UTC (7,939 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models, by Yidi Wang and 6 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.RO

new | recent | 2026-09

Change to browse by:

cs

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • R2S-Eval combines real-to-sim calibration with VLM-based preference evaluation to move beyond manual success counting.
  • Simulator-generated rollout videos are calibrated to real-world evaluation settings, reducing repeated hardware trials and manual scene resets.
  • A VLM evaluator performs pairwise quality comparisons that are aggregated into stable, human-aligned policy rankings.
  • Experiments in simulation and real-world settings show behavior-quality differences that binary success metrics miss.

Highlights and analysis are generated automatically and may contain errors. Check the original source.