Evaluating Gemini Robotics Policies in a Veo World Simulator

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gemini Robotics Team, Choromanski, Krzysztof, Devin, Coline, Du, Yilun, Dwibedi, Debidatta, Gao, Ruiqi, Jindal, Abhishek, Kipf, Thomas, Kirmani, Sean, Leal, Isabel, Liu, Fangchen, Majumdar, Anirudha, Marmon, Andrew, Parada, Carolina, Rubanova, Yulia, Shah, Dhruv, Sindhwani, Vikas, Tan, Jie, Xia, Fei, Xiao, Ted, Yang, Sherry, Yu, Wenhao, Zhou, Allan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909982533877760
author Gemini Robotics Team
Choromanski, Krzysztof
Devin, Coline
Du, Yilun
Dwibedi, Debidatta
Gao, Ruiqi
Jindal, Abhishek
Kipf, Thomas
Kirmani, Sean
Leal, Isabel
Liu, Fangchen
Majumdar, Anirudha
Marmon, Andrew
Parada, Carolina
Rubanova, Yulia
Shah, Dhruv
Sindhwani, Vikas
Tan, Jie
Xia, Fei
Xiao, Ted
Yang, Sherry
Yu, Wenhao
Zhou, Allan
author_facet Gemini Robotics Team
Choromanski, Krzysztof
Devin, Coline
Du, Yilun
Dwibedi, Debidatta
Gao, Ruiqi
Jindal, Abhishek
Kipf, Thomas
Kirmani, Sean
Leal, Isabel
Liu, Fangchen
Majumdar, Anirudha
Marmon, Andrew
Parada, Carolina
Rubanova, Yulia
Shah, Dhruv
Sindhwani, Vikas
Tan, Jie
Xia, Fei
Xiao, Ted
Yang, Sherry
Yu, Wenhao
Zhou, Allan
contents Generative world models hold significant potential for simulating interactions with visuomotor policies in varied environments. Frontier video models can enable generation of realistic observations and environment interactions in a scalable and general manner. However, the use of video models in robotics has been limited primarily to in-distribution evaluations, i.e., scenarios that are similar to ones used to train the policy or fine-tune the base video model. In this report, we demonstrate that video models can be used for the entire spectrum of policy evaluation use cases in robotics: from assessing nominal performance to out-of-distribution (OOD) generalization, and probing physical and semantic safety. We introduce a generative evaluation system built upon a frontier video foundation model (Veo). The system is optimized to support robot action conditioning and multi-view consistency, while integrating generative image-editing and multi-view completion to synthesize realistic variations of real-world scenes along multiple axes of generalization. We demonstrate that the system preserves the base capabilities of the video model to enable accurate simulation of scenes that have been edited to include novel interaction objects, novel visual backgrounds, and novel distractor objects. This fidelity enables accurately predicting the relative performance of different policies in both nominal and OOD conditions, determining the relative impact of different axes of generalization on policy performance, and performing red teaming of policies to expose behaviors that violate physical or semantic safety constraints. We validate these capabilities through 1600+ real-world evaluations of eight Gemini Robotics policy checkpoints and five tasks for a bimanual manipulator.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10675
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Gemini Robotics Policies in a Veo World Simulator
Gemini Robotics Team
Choromanski, Krzysztof
Devin, Coline
Du, Yilun
Dwibedi, Debidatta
Gao, Ruiqi
Jindal, Abhishek
Kipf, Thomas
Kirmani, Sean
Leal, Isabel
Liu, Fangchen
Majumdar, Anirudha
Marmon, Andrew
Parada, Carolina
Rubanova, Yulia
Shah, Dhruv
Sindhwani, Vikas
Tan, Jie
Xia, Fei
Xiao, Ted
Yang, Sherry
Yu, Wenhao
Zhou, Allan
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Generative world models hold significant potential for simulating interactions with visuomotor policies in varied environments. Frontier video models can enable generation of realistic observations and environment interactions in a scalable and general manner. However, the use of video models in robotics has been limited primarily to in-distribution evaluations, i.e., scenarios that are similar to ones used to train the policy or fine-tune the base video model. In this report, we demonstrate that video models can be used for the entire spectrum of policy evaluation use cases in robotics: from assessing nominal performance to out-of-distribution (OOD) generalization, and probing physical and semantic safety. We introduce a generative evaluation system built upon a frontier video foundation model (Veo). The system is optimized to support robot action conditioning and multi-view consistency, while integrating generative image-editing and multi-view completion to synthesize realistic variations of real-world scenes along multiple axes of generalization. We demonstrate that the system preserves the base capabilities of the video model to enable accurate simulation of scenes that have been edited to include novel interaction objects, novel visual backgrounds, and novel distractor objects. This fidelity enables accurately predicting the relative performance of different policies in both nominal and OOD conditions, determining the relative impact of different axes of generalization on policy performance, and performing red teaming of policies to expose behaviors that violate physical or semantic safety constraints. We validate these capabilities through 1600+ real-world evaluations of eight Gemini Robotics policy checkpoints and five tasks for a bimanual manipulator.
title Evaluating Gemini Robotics Policies in a Veo World Simulator
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2512.10675