Robot Data Curation with Mutual Information Estimators

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hejna, Joey, Mirchandani, Suvir, Balakrishna, Ashwin, Xie, Annie, Wahid, Ayzaan, Tompson, Jonathan, Sanketi, Pannag, Shah, Dhruv, Devin, Coline, Sadigh, Dorsa
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910916027613184
author Hejna, Joey
Mirchandani, Suvir
Balakrishna, Ashwin
Xie, Annie
Wahid, Ayzaan
Tompson, Jonathan
Sanketi, Pannag
Shah, Dhruv
Devin, Coline
Sadigh, Dorsa
author_facet Hejna, Joey
Mirchandani, Suvir
Balakrishna, Ashwin
Xie, Annie
Wahid, Ayzaan
Tompson, Jonathan
Sanketi, Pannag
Shah, Dhruv
Devin, Coline
Sadigh, Dorsa
contents The performance of imitation learning policies often hinges on the datasets with which they are trained. Consequently, investment in data collection for robotics has grown across both industrial and academic labs. However, despite the marked increase in the quantity of demonstrations collected, little work has sought to assess the quality of said data despite mounting evidence of its importance in other areas such as vision and language. In this work, we take a critical step towards addressing the data quality in robotics. Given a dataset of demonstrations, we aim to estimate the relative quality of individual demonstrations in terms of both action diversity and predictability. To do so, we estimate the average contribution of a trajectory towards the mutual information between states and actions in the entire dataset, which captures both the entropy of the marginal action distribution and the state-conditioned action entropy. Though commonly used mutual information estimators require vast amounts of data often beyond the scale available in robotics, we introduce a novel technique based on k-nearest neighbor estimates of mutual information on top of simple VAE embeddings of states and actions. Empirically, we demonstrate that our approach is able to partition demonstration datasets by quality according to human expert scores across a diverse set of benchmarks spanning simulation and real world environments. Moreover, training policies based on data filtered by our method leads to a 5-10% improvement in RoboMimic and better performance on real ALOHA and Franka setups.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robot Data Curation with Mutual Information Estimators
Hejna, Joey
Mirchandani, Suvir
Balakrishna, Ashwin
Xie, Annie
Wahid, Ayzaan
Tompson, Jonathan
Sanketi, Pannag
Shah, Dhruv
Devin, Coline
Sadigh, Dorsa
Robotics
The performance of imitation learning policies often hinges on the datasets with which they are trained. Consequently, investment in data collection for robotics has grown across both industrial and academic labs. However, despite the marked increase in the quantity of demonstrations collected, little work has sought to assess the quality of said data despite mounting evidence of its importance in other areas such as vision and language. In this work, we take a critical step towards addressing the data quality in robotics. Given a dataset of demonstrations, we aim to estimate the relative quality of individual demonstrations in terms of both action diversity and predictability. To do so, we estimate the average contribution of a trajectory towards the mutual information between states and actions in the entire dataset, which captures both the entropy of the marginal action distribution and the state-conditioned action entropy. Though commonly used mutual information estimators require vast amounts of data often beyond the scale available in robotics, we introduce a novel technique based on k-nearest neighbor estimates of mutual information on top of simple VAE embeddings of states and actions. Empirically, we demonstrate that our approach is able to partition demonstration datasets by quality according to human expert scores across a diverse set of benchmarks spanning simulation and real world environments. Moreover, training policies based on data filtered by our method leads to a 5-10% improvement in RoboMimic and better performance on real ALOHA and Franka setups.
title Robot Data Curation with Mutual Information Estimators
topic Robotics
url https://arxiv.org/abs/2502.08623