R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Albert, Huang, Tzu-Heng, Cooper, John, Trost, Avi, Chu, Ziyi, GNVV, Satya Sai Srinath Namburi, Cai, Ziyang, Park, Kendall, Roberts, Nicholas, Sala, Frederic
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910924466552832
author Ge, Albert
Huang, Tzu-Heng
Cooper, John
Trost, Avi
Chu, Ziyi
GNVV, Satya Sai Srinath Namburi
Cai, Ziyang
Park, Kendall
Roberts, Nicholas
Sala, Frederic
author_facet Ge, Albert
Huang, Tzu-Heng
Cooper, John
Trost, Avi
Chu, Ziyi
GNVV, Satya Sai Srinath Namburi
Cai, Ziyang
Park, Kendall
Roberts, Nicholas
Sala, Frederic
contents Data mixing strategies have successfully reduced the costs involved in training language models. While promising, such methods suffer from two flaws. First, they rely on predetermined data domains (e.g., data sources, task types), which may fail to capture critical semantic nuances, leaving performance on the table. Second, these methods scale with the number of domains in a computationally prohibitive way. We address these challenges via R&B, a framework that re-partitions training data based on semantic similarity (Regroup) to create finer-grained domains, and efficiently optimizes the data composition (Balance) by leveraging a Gram matrix induced by domain gradients obtained throughout training. Unlike prior works, it removes the need for additional compute to obtain evaluation information such as losses or gradients. We analyze this technique under standard regularity conditions and provide theoretical insights that justify R&B's effectiveness compared to non-adaptive mixing approaches. Empirically, we demonstrate the effectiveness of R&B on five diverse datasets ranging from natural language to reasoning and multimodal tasks. With as little as 0.01% additional compute overhead, R&B matches or exceeds the performance of state-of-the-art data mixing strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00358
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
Ge, Albert
Huang, Tzu-Heng
Cooper, John
Trost, Avi
Chu, Ziyi
GNVV, Satya Sai Srinath Namburi
Cai, Ziyang
Park, Kendall
Roberts, Nicholas
Sala, Frederic
Machine Learning
Artificial Intelligence
Computation and Language
Data mixing strategies have successfully reduced the costs involved in training language models. While promising, such methods suffer from two flaws. First, they rely on predetermined data domains (e.g., data sources, task types), which may fail to capture critical semantic nuances, leaving performance on the table. Second, these methods scale with the number of domains in a computationally prohibitive way. We address these challenges via R&B, a framework that re-partitions training data based on semantic similarity (Regroup) to create finer-grained domains, and efficiently optimizes the data composition (Balance) by leveraging a Gram matrix induced by domain gradients obtained throughout training. Unlike prior works, it removes the need for additional compute to obtain evaluation information such as losses or gradients. We analyze this technique under standard regularity conditions and provide theoretical insights that justify R&B's effectiveness compared to non-adaptive mixing approaches. Empirically, we demonstrate the effectiveness of R&B on five diverse datasets ranging from natural language to reasoning and multimodal tasks. With as little as 0.01% additional compute overhead, R&B matches or exceeds the performance of state-of-the-art data mixing strategies.
title R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.00358