Guiding Data Collection via Factored Scaling Curves

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zha, Lihan, Badithela, Apurva, Zhang, Michael, Lidard, Justin, Bao, Jeremy, Zhou, Emily, Snyder, David, Ren, Allen Z., Shah, Dhruv, Majumdar, Anirudha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909607834681344
author Zha, Lihan
Badithela, Apurva
Zhang, Michael
Lidard, Justin
Bao, Jeremy
Zhou, Emily
Snyder, David
Ren, Allen Z.
Shah, Dhruv
Majumdar, Anirudha
author_facet Zha, Lihan
Badithela, Apurva
Zhang, Michael
Lidard, Justin
Bao, Jeremy
Zhou, Emily
Snyder, David
Ren, Allen Z.
Shah, Dhruv
Majumdar, Anirudha
contents Generalist imitation learning policies trained on large datasets show great promise for solving diverse manipulation tasks. However, to ensure generalization to different conditions, policies need to be trained with data collected across a large set of environmental factor variations (e.g., camera pose, table height, distractors) $-$ a prohibitively expensive undertaking, if done exhaustively. We introduce a principled method for deciding what data to collect and how much to collect for each factor by constructing factored scaling curves (FSC), which quantify how policy performance varies as data scales along individual or paired factors. These curves enable targeted data acquisition for the most influential factor combinations within a given budget. We evaluate the proposed method through extensive simulated and real-world experiments, across both training-from-scratch and fine-tuning settings, and show that it boosts success rates in real-world tasks in new environments by up to 26% over existing data-collection strategies. We further demonstrate how factored scaling curves can effectively guide data collection using an offline metric, without requiring real-world evaluation at scale.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07728
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Guiding Data Collection via Factored Scaling Curves
Zha, Lihan
Badithela, Apurva
Zhang, Michael
Lidard, Justin
Bao, Jeremy
Zhou, Emily
Snyder, David
Ren, Allen Z.
Shah, Dhruv
Majumdar, Anirudha
Robotics
Artificial Intelligence
Machine Learning
Generalist imitation learning policies trained on large datasets show great promise for solving diverse manipulation tasks. However, to ensure generalization to different conditions, policies need to be trained with data collected across a large set of environmental factor variations (e.g., camera pose, table height, distractors) $-$ a prohibitively expensive undertaking, if done exhaustively. We introduce a principled method for deciding what data to collect and how much to collect for each factor by constructing factored scaling curves (FSC), which quantify how policy performance varies as data scales along individual or paired factors. These curves enable targeted data acquisition for the most influential factor combinations within a given budget. We evaluate the proposed method through extensive simulated and real-world experiments, across both training-from-scratch and fine-tuning settings, and show that it boosts success rates in real-world tasks in new environments by up to 26% over existing data-collection strategies. We further demonstrate how factored scaling curves can effectively guide data collection using an offline metric, without requiring real-world evaluation at scale.
title Guiding Data Collection via Factored Scaling Curves
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.07728