Saved in:
Bibliographic Details
Main Authors: Wang, Haohui, Qi, Jingyuan, Chen, Jianpeng, Wu, Jun, Huang, Lifu, Zheng, Lecheng, Choi, Kevin, Veeramani, Balaji, Bowen, Edward, Hu, Alison, Cody, Tyler, Zhou, Dawei
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.13640
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918206312022016
author Wang, Haohui
Qi, Jingyuan
Chen, Jianpeng
Wu, Jun
Huang, Lifu
Zheng, Lecheng
Choi, Kevin
Veeramani, Balaji
Bowen, Edward
Hu, Alison
Cody, Tyler
Zhou, Dawei
author_facet Wang, Haohui
Qi, Jingyuan
Chen, Jianpeng
Wu, Jun
Huang, Lifu
Zheng, Lecheng
Choi, Kevin
Veeramani, Balaji
Bowen, Edward
Hu, Alison
Cody, Tyler
Zhou, Dawei
contents The rapid progress of large language models (LLMs) is fueled by the growing reliance on datasets that blend real and synthetic data. While synthetic data offers scalability and cost-efficiency, it often introduces systematic distributional discrepancies, particularly underrepresenting long-tail knowledge due to truncation effects from data generation mechanisms like top-p sampling, temperature scaling, and finite sampling. These discrepancies pose fundamental challenges in characterizing and evaluating the utility of mixed real-synthetic datasets. In this paper, we identify a three-phase scaling behavior characterized by two breakpoints that reflect transitions in model behavior across learning head and tail knowledge. We further derive an LLM generalization bound designed for real and synthetic mixtures, revealing several key factors that govern their generalization performance. Building on our theoretical findings, we propose an effective yet efficient data valuation method that scales to large-scale datasets. Comprehensive experiments across four tasks, including image classification, sentiment classification, instruction following, and complex reasoning, demonstrate that our method surpasses state-of-the-art baselines in data valuation with significantly low computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13640
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
Wang, Haohui
Qi, Jingyuan
Chen, Jianpeng
Wu, Jun
Huang, Lifu
Zheng, Lecheng
Choi, Kevin
Veeramani, Balaji
Bowen, Edward
Hu, Alison
Cody, Tyler
Zhou, Dawei
Machine Learning
Artificial Intelligence
The rapid progress of large language models (LLMs) is fueled by the growing reliance on datasets that blend real and synthetic data. While synthetic data offers scalability and cost-efficiency, it often introduces systematic distributional discrepancies, particularly underrepresenting long-tail knowledge due to truncation effects from data generation mechanisms like top-p sampling, temperature scaling, and finite sampling. These discrepancies pose fundamental challenges in characterizing and evaluating the utility of mixed real-synthetic datasets. In this paper, we identify a three-phase scaling behavior characterized by two breakpoints that reflect transitions in model behavior across learning head and tail knowledge. We further derive an LLM generalization bound designed for real and synthetic mixtures, revealing several key factors that govern their generalization performance. Building on our theoretical findings, we propose an effective yet efficient data valuation method that scales to large-scale datasets. Comprehensive experiments across four tasks, including image classification, sentiment classification, instruction following, and complex reasoning, demonstrate that our method surpasses state-of-the-art baselines in data valuation with significantly low computational cost.
title Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.13640