Using Temperature Sampling to Effectively Train Robot Learning Policies on Imbalanced Datasets

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Patil, Basavasagar, Belt, Sydney, Lee, Jayjun, Fazeli, Nima, Bucher, Bernadette
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908605961207808
author Patil, Basavasagar
Belt, Sydney
Lee, Jayjun
Fazeli, Nima
Bucher, Bernadette
author_facet Patil, Basavasagar
Belt, Sydney
Lee, Jayjun
Fazeli, Nima
Bucher, Bernadette
contents Increasingly large datasets of robot actions and sensory observations are being collected to train ever-larger neural networks. These datasets are collected based on tasks and while these tasks may be distinct in their descriptions, many involve very similar physical action sequences (e.g., 'pick up an apple' versus 'pick up an orange'). As a result, many datasets of robotic tasks are substantially imbalanced in terms of the physical robotic actions they represent. In this work, we propose a simple sampling strategy for policy training that mitigates this imbalance. Our method requires only a few lines of code to integrate into existing codebases and improves generalization. We evaluate our method in both pre-training small models and fine-tuning large foundational models. Our results show substantial improvements on low-resource tasks compared to prior state-of-the-art methods, without degrading performance on high-resource tasks. This enables more effective use of model capacity for multi-task policies. We also further validate our approach in a real-world setup on a Franka Panda robot arm across a diverse set of tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Using Temperature Sampling to Effectively Train Robot Learning Policies on Imbalanced Datasets
Patil, Basavasagar
Belt, Sydney
Lee, Jayjun
Fazeli, Nima
Bucher, Bernadette
Robotics
Machine Learning
Increasingly large datasets of robot actions and sensory observations are being collected to train ever-larger neural networks. These datasets are collected based on tasks and while these tasks may be distinct in their descriptions, many involve very similar physical action sequences (e.g., 'pick up an apple' versus 'pick up an orange'). As a result, many datasets of robotic tasks are substantially imbalanced in terms of the physical robotic actions they represent. In this work, we propose a simple sampling strategy for policy training that mitigates this imbalance. Our method requires only a few lines of code to integrate into existing codebases and improves generalization. We evaluate our method in both pre-training small models and fine-tuning large foundational models. Our results show substantial improvements on low-resource tasks compared to prior state-of-the-art methods, without degrading performance on high-resource tasks. This enables more effective use of model capacity for multi-task policies. We also further validate our approach in a real-world setup on a Franka Panda robot arm across a diverse set of tasks.
title Using Temperature Sampling to Effectively Train Robot Learning Policies on Imbalanced Datasets
topic Robotics
Machine Learning
url https://arxiv.org/abs/2510.19373