Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916990930649088 |
|---|---|
| author | Badithela, Apurva Snyder, David Zha, Lihan Mikhail, Joseph O'Kelly, Matthew Dixit, Anushri Majumdar, Anirudha |
| author_facet | Badithela, Apurva Snyder, David Zha, Lihan Mikhail, Joseph O'Kelly, Matthew Dixit, Anushri Majumdar, Anirudha |
| contents | Rapid progress in imitation learning, foundation models, and large-scale datasets has led to robot manipulation policies that generalize to a wide-range of tasks and environments. However, rigorous evaluation of these policies remains a challenge. Typically in practice, robot policies are often evaluated on a small number of hardware trials without any statistical assurances. We present SureSim, a framework to augment large-scale simulation with relatively small-scale real-world testing to provide reliable inferences on the real-world performance of a policy. Our key idea is to formalize the problem of combining real and simulation evaluations as a prediction-powered inference problem, in which a small number of paired real and simulation evaluations are used to rectify bias in large-scale simulation. We then leverage non-asymptotic mean estimation algorithms to provide confidence intervals on mean policy performance. Using physics-based simulation, we evaluate both diffusion policy and multi-task fine-tuned \(π_0\) on a joint distribution of objects and initial conditions, and find that our approach saves over \(20-25\%\) of hardware evaluation effort to achieve similar bounds on policy performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_04354 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators Badithela, Apurva Snyder, David Zha, Lihan Mikhail, Joseph O'Kelly, Matthew Dixit, Anushri Majumdar, Anirudha Robotics Artificial Intelligence Systems and Control Rapid progress in imitation learning, foundation models, and large-scale datasets has led to robot manipulation policies that generalize to a wide-range of tasks and environments. However, rigorous evaluation of these policies remains a challenge. Typically in practice, robot policies are often evaluated on a small number of hardware trials without any statistical assurances. We present SureSim, a framework to augment large-scale simulation with relatively small-scale real-world testing to provide reliable inferences on the real-world performance of a policy. Our key idea is to formalize the problem of combining real and simulation evaluations as a prediction-powered inference problem, in which a small number of paired real and simulation evaluations are used to rectify bias in large-scale simulation. We then leverage non-asymptotic mean estimation algorithms to provide confidence intervals on mean policy performance. Using physics-based simulation, we evaluate both diffusion policy and multi-task fine-tuned \(π_0\) on a joint distribution of objects and initial conditions, and find that our approach saves over \(20-25\%\) of hardware evaluation effort to achieve similar bounds on policy performance. |
| title | Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators |
| topic | Robotics Artificial Intelligence Systems and Control |
| url | https://arxiv.org/abs/2510.04354 |