OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Budzianowski, Paweł, Wiśnios, Emilia, Tyrolski, Michał, Góral, Gracjan, Kulakov, Igor, Petrenko, Viktor, Walas, Krzysztof
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910015997083648
author Budzianowski, Paweł
Wiśnios, Emilia
Tyrolski, Michał
Góral, Gracjan
Kulakov, Igor
Petrenko, Viktor
Walas, Krzysztof
author_facet Budzianowski, Paweł
Wiśnios, Emilia
Tyrolski, Michał
Góral, Gracjan
Kulakov, Igor
Petrenko, Viktor
Walas, Krzysztof
contents Data scarcity remains one of the most limiting factors in driving progress in robotics. However, the amount of available robotics data in the wild is growing exponentially, creating new opportunities for large-scale data utilization. Reliable temporal task completion prediction could help automatically annotate and curate this data at scale. The Generative Value Learning (GVL) approach was recently proposed, leveraging the knowledge embedded in vision-language models (VLMs) to predict task progress from visual observations. Building upon GVL, we propose OpenGVL, a comprehensive benchmark for estimating task progress across diverse challenging manipulation tasks involving both robotic and human embodiments. We evaluate the capabilities of publicly available open-source foundation models, showing that open-source model families significantly underperform closed-source counterparts, achieving only approximately $70\%$ of their performance on temporal progress prediction tasks. Furthermore, we demonstrate how OpenGVL can serve as a practical tool for automated data curation and filtering, enabling efficient quality assessment of large-scale robotics datasets. We release the benchmark along with the complete codebase at \href{github.com/budzianowski/opengvl}{OpenGVL}.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17321
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
Budzianowski, Paweł
Wiśnios, Emilia
Tyrolski, Michał
Góral, Gracjan
Kulakov, Igor
Petrenko, Viktor
Walas, Krzysztof
Robotics
Computation and Language
Data scarcity remains one of the most limiting factors in driving progress in robotics. However, the amount of available robotics data in the wild is growing exponentially, creating new opportunities for large-scale data utilization. Reliable temporal task completion prediction could help automatically annotate and curate this data at scale. The Generative Value Learning (GVL) approach was recently proposed, leveraging the knowledge embedded in vision-language models (VLMs) to predict task progress from visual observations. Building upon GVL, we propose OpenGVL, a comprehensive benchmark for estimating task progress across diverse challenging manipulation tasks involving both robotic and human embodiments. We evaluate the capabilities of publicly available open-source foundation models, showing that open-source model families significantly underperform closed-source counterparts, achieving only approximately $70\%$ of their performance on temporal progress prediction tasks. Furthermore, we demonstrate how OpenGVL can serve as a practical tool for automated data curation and filtering, enabling efficient quality assessment of large-scale robotics datasets. We release the benchmark along with the complete codebase at \href{github.com/budzianowski/opengvl}{OpenGVL}.
title OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
topic Robotics
Computation and Language
url https://arxiv.org/abs/2509.17321