REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909973677604864 |
|---|---|
| author | Sedlacek, Martin Yefanov, Pavlo Ponimatkin, Georgy Bardhan, Jai Pilc, Simon Fourmy, Mederic Kazakos, Evangelos Snoek, Cees G. M. Sivic, Josef Petrik, Vladimir |
| author_facet | Sedlacek, Martin Yefanov, Pavlo Ponimatkin, Georgy Bardhan, Jai Pilc, Simon Fourmy, Mederic Kazakos, Evangelos Snoek, Cees G. M. Sivic, Josef Petrik, Vladimir |
| contents | Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to generalize beyond the specific environments and conditions they were trained on, which is presently difficult and expensive to evaluate in the real-world. To address this gap, we present REALM, a new simulation environment and benchmark designed to evaluate the generalization capabilities of VLA models, with a specific emphasis on establishing a strong correlation between simulated and real-world performance through high-fidelity visuals and aligned robot control. Our environment offers a suite of 15 perturbation factors, 7 manipulation skills, and more than 3,500 objects. Finally, we establish two task sets that form our benchmark and evaluate the π_{0}, π_{0}-FAST, and GR00T N1.5 VLA models, showing that generalization and robustness remain an open challenge. More broadly, we also show that simulation gives us a valuable proxy for the real-world and allows us to systematically probe for and quantify the weaknesses and failure modes of VLAs. Project page: https://martin-sedlacek.com/realm |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_19562 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation Sedlacek, Martin Yefanov, Pavlo Ponimatkin, Georgy Bardhan, Jai Pilc, Simon Fourmy, Mederic Kazakos, Evangelos Snoek, Cees G. M. Sivic, Josef Petrik, Vladimir Robotics Artificial Intelligence Vision-Language-Action (VLA) models empower robots to understand and execute tasks described by natural language instructions. However, a key challenge lies in their ability to generalize beyond the specific environments and conditions they were trained on, which is presently difficult and expensive to evaluate in the real-world. To address this gap, we present REALM, a new simulation environment and benchmark designed to evaluate the generalization capabilities of VLA models, with a specific emphasis on establishing a strong correlation between simulated and real-world performance through high-fidelity visuals and aligned robot control. Our environment offers a suite of 15 perturbation factors, 7 manipulation skills, and more than 3,500 objects. Finally, we establish two task sets that form our benchmark and evaluate the π_{0}, π_{0}-FAST, and GR00T N1.5 VLA models, showing that generalization and robustness remain an open challenge. More broadly, we also show that simulation gives us a valuable proxy for the real-world and allows us to systematically probe for and quantify the weaknesses and failure modes of VLAs. Project page: https://martin-sedlacek.com/realm |
| title | REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation |
| topic | Robotics Artificial Intelligence |
| url | https://arxiv.org/abs/2512.19562 |