Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Badithela, Apurva, Snyder, David, Zha, Lihan, Mikhail, Joseph, O'Kelly, Matthew, Dixit, Anushri, Majumdar, Anirudha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916990930649088
author Badithela, Apurva
Snyder, David
Zha, Lihan
Mikhail, Joseph
O'Kelly, Matthew
Dixit, Anushri
Majumdar, Anirudha
author_facet Badithela, Apurva
Snyder, David
Zha, Lihan
Mikhail, Joseph
O'Kelly, Matthew
Dixit, Anushri
Majumdar, Anirudha
contents Rapid progress in imitation learning, foundation models, and large-scale datasets has led to robot manipulation policies that generalize to a wide-range of tasks and environments. However, rigorous evaluation of these policies remains a challenge. Typically in practice, robot policies are often evaluated on a small number of hardware trials without any statistical assurances. We present SureSim, a framework to augment large-scale simulation with relatively small-scale real-world testing to provide reliable inferences on the real-world performance of a policy. Our key idea is to formalize the problem of combining real and simulation evaluations as a prediction-powered inference problem, in which a small number of paired real and simulation evaluations are used to rectify bias in large-scale simulation. We then leverage non-asymptotic mean estimation algorithms to provide confidence intervals on mean policy performance. Using physics-based simulation, we evaluate both diffusion policy and multi-task fine-tuned \(π_0\) on a joint distribution of objects and initial conditions, and find that our approach saves over \(20-25\%\) of hardware evaluation effort to achieve similar bounds on policy performance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators
Badithela, Apurva
Snyder, David
Zha, Lihan
Mikhail, Joseph
O'Kelly, Matthew
Dixit, Anushri
Majumdar, Anirudha
Robotics
Artificial Intelligence
Systems and Control
Rapid progress in imitation learning, foundation models, and large-scale datasets has led to robot manipulation policies that generalize to a wide-range of tasks and environments. However, rigorous evaluation of these policies remains a challenge. Typically in practice, robot policies are often evaluated on a small number of hardware trials without any statistical assurances. We present SureSim, a framework to augment large-scale simulation with relatively small-scale real-world testing to provide reliable inferences on the real-world performance of a policy. Our key idea is to formalize the problem of combining real and simulation evaluations as a prediction-powered inference problem, in which a small number of paired real and simulation evaluations are used to rectify bias in large-scale simulation. We then leverage non-asymptotic mean estimation algorithms to provide confidence intervals on mean policy performance. Using physics-based simulation, we evaluate both diffusion policy and multi-task fine-tuned \(π_0\) on a joint distribution of objects and initial conditions, and find that our approach saves over \(20-25\%\) of hardware evaluation effort to achieve similar bounds on policy performance.
title Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators
topic Robotics
Artificial Intelligence
Systems and Control
url https://arxiv.org/abs/2510.04354