Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Haohui, Qi, Jingyuan, Chen, Jianpeng, Wu, Jun, Huang, Lifu, Zheng, Lecheng, Choi, Kevin, Veeramani, Balaji, Bowen, Edward, Hu, Alison, Cody, Tyler, Zhou, Dawei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EvoluNet: Advancing Dynamic Non-IID Transfer Learning on Graphs
di: Wang, Haohui, et al.
Pubblicazione: (2023)
di: Wang, Haohui, et al.
Pubblicazione: (2023)
LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection
di: Zeng, Xinyue, et al.
Pubblicazione: (2025)
di: Zeng, Xinyue, et al.
Pubblicazione: (2025)
MetamatBench: Integrating Heterogeneous Data, Computational Tools, and Visual Interface for Metamaterial Discovery
di: Chen, Jianpeng, et al.
Pubblicazione: (2025)
di: Chen, Jianpeng, et al.
Pubblicazione: (2025)
Scaling Laws for Mixture Pretraining Under Data Constraints
di: Sedova, Anastasiia, et al.
Pubblicazione: (2026)
di: Sedova, Anastasiia, et al.
Pubblicazione: (2026)
Towards Heterogeneous Long-tailed Learning: Benchmarking, Metrics, and Toolbox
di: Wang, Haohui, et al.
Pubblicazione: (2023)
di: Wang, Haohui, et al.
Pubblicazione: (2023)
Indiaʼs Integration Into Global Value Chains: Policy Approaches to Growth and Industrial Upgrading
di: Veeramani Choorikkadan
Pubblicazione: (2026)
di: Veeramani Choorikkadan
Pubblicazione: (2026)
A Framework for Assessing the Value of RFID Implementation by Tier-One Suppliers to Major Retailers
di: Dharmaraj Veeramani
Pubblicazione: (2008)
di: Dharmaraj Veeramani
Pubblicazione: (2008)
Principled Synthetic Data Enables the First Scaling Laws for LLMs in Recommendation
di: Zhang, Benyu, et al.
Pubblicazione: (2026)
di: Zhang, Benyu, et al.
Pubblicazione: (2026)
Beyond Real Data: Synthetic Data through the Lens of Regularization
di: Shidani, Amitis, et al.
Pubblicazione: (2025)
di: Shidani, Amitis, et al.
Pubblicazione: (2025)
Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
di: Chen, Qi, et al.
Pubblicazione: (2025)
di: Chen, Qi, et al.
Pubblicazione: (2025)
Scaling Laws for Optimal Data Mixtures
di: Shukor, Mustafa, et al.
Pubblicazione: (2025)
di: Shukor, Mustafa, et al.
Pubblicazione: (2025)
A Large-Scale Real-World Evaluation of LLM-Based Virtual Teaching Assistant
di: Kweon, Sunjun, et al.
Pubblicazione: (2025)
di: Kweon, Sunjun, et al.
Pubblicazione: (2025)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
di: Kim, Taehyun, et al.
Pubblicazione: (2024)
di: Kim, Taehyun, et al.
Pubblicazione: (2024)
Realistic Synthetic Household Data Generation at Scale
di: Singh, Siddharth, et al.
Pubblicazione: (2026)
di: Singh, Siddharth, et al.
Pubblicazione: (2026)
Scaling Laws of Synthetic Data for Language Models
di: Qin, Zeyu, et al.
Pubblicazione: (2025)
di: Qin, Zeyu, et al.
Pubblicazione: (2025)
Scaling Up Forest Vision with Synthetic Data
di: She, Yihang, et al.
Pubblicazione: (2025)
di: She, Yihang, et al.
Pubblicazione: (2025)
Patient-Zero: Scaling Synthetic Patient Agents to Real-World Distributions without Real Patient Data
di: Lai, Yunghwei, et al.
Pubblicazione: (2025)
di: Lai, Yunghwei, et al.
Pubblicazione: (2025)
DKT2: Revisiting Applicable and Comprehensive Knowledge Tracing in Large-Scale Data
di: Zhou, Yiyun, et al.
Pubblicazione: (2025)
di: Zhou, Yiyun, et al.
Pubblicazione: (2025)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
di: Pang, Tianyu, et al.
Pubblicazione: (2026)
di: Pang, Tianyu, et al.
Pubblicazione: (2026)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
di: Yang, Yue, et al.
Pubblicazione: (2025)
di: Yang, Yue, et al.
Pubblicazione: (2025)
Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training
di: Ruan, Hongzhi, et al.
Pubblicazione: (2026)
di: Ruan, Hongzhi, et al.
Pubblicazione: (2026)
Characterizing Model Behavior Under Synthetic Data Training: An Empirical Study Across Scales and Mixing Ratios
di: Du, Y., et al.
Pubblicazione: (2025)
di: Du, Y., et al.
Pubblicazione: (2025)
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
di: Askari-Hemmat, Reyhane, et al.
Pubblicazione: (2025)
di: Askari-Hemmat, Reyhane, et al.
Pubblicazione: (2025)
Synthetica: Large Scale Synthetic Data for Robot Perception
di: Singh, Ritvik, et al.
Pubblicazione: (2024)
di: Singh, Ritvik, et al.
Pubblicazione: (2024)
Scaling Trends for Data Poisoning in LLMs
di: Bowen, Dillon, et al.
Pubblicazione: (2024)
di: Bowen, Dillon, et al.
Pubblicazione: (2024)
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
di: DatologyAI, et al.
Pubblicazione: (2025)
di: DatologyAI, et al.
Pubblicazione: (2025)
SynPlay: Large-Scale Synthetic Human Data with Real-World Diversity for Aerial-View Perception
di: Yim, Jinsub, et al.
Pubblicazione: (2024)
di: Yim, Jinsub, et al.
Pubblicazione: (2024)
Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
di: Pan, Zhixuan, et al.
Pubblicazione: (2025)
di: Pan, Zhixuan, et al.
Pubblicazione: (2025)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
di: Borisov, Vadim
Pubblicazione: (2026)
di: Borisov, Vadim
Pubblicazione: (2026)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
Scaling Planning for Automated Driving using Simplistic Synthetic Data
di: Stoll, Martin, et al.
Pubblicazione: (2023)
di: Stoll, Martin, et al.
Pubblicazione: (2023)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
DALDA: Data Augmentation Leveraging Diffusion Model and LLM with Adaptive Guidance Scaling
di: Jung, Kyuheon, et al.
Pubblicazione: (2024)
di: Jung, Kyuheon, et al.
Pubblicazione: (2024)
Addressing Bias in Spoken Language Systems Used in the Development and Implementation of Automated Child Language‐Based Assessment
di: Alison L. Bailey, et al.
Pubblicazione: (2025)
di: Alison L. Bailey, et al.
Pubblicazione: (2025)
Citations and Trust in LLM Generated Responses
di: Ding, Yifan, et al.
Pubblicazione: (2025)
di: Ding, Yifan, et al.
Pubblicazione: (2025)
Triply Laplacian Scale Mixture Modeling for Seismic Data Noise Suppression
di: Pan, Sirui, et al.
Pubblicazione: (2025)
di: Pan, Sirui, et al.
Pubblicazione: (2025)
Re-Mix: Optimizing Data Mixtures for Large Scale Imitation Learning
di: Hejna, Joey, et al.
Pubblicazione: (2024)
di: Hejna, Joey, et al.
Pubblicazione: (2024)
Large-Scale Multiple Testing of Composite Null Hypotheses Under Heteroskedasticity
di: Gang, Bowen, et al.
Pubblicazione: (2023)
di: Gang, Bowen, et al.
Pubblicazione: (2023)
On Data Engineering for Scaling LLM Terminal Capabilities
di: Pi, Renjie, et al.
Pubblicazione: (2026)
di: Pi, Renjie, et al.
Pubblicazione: (2026)
Documenti analoghi
-
EvoluNet: Advancing Dynamic Non-IID Transfer Learning on Graphs
di: Wang, Haohui, et al.
Pubblicazione: (2023) -
LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection
di: Zeng, Xinyue, et al.
Pubblicazione: (2025) -
MetamatBench: Integrating Heterogeneous Data, Computational Tools, and Visual Interface for Metamaterial Discovery
di: Chen, Jianpeng, et al.
Pubblicazione: (2025) -
Scaling Laws for Mixture Pretraining Under Data Constraints
di: Sedova, Anastasiia, et al.
Pubblicazione: (2026) -
Towards Heterogeneous Long-tailed Learning: Benchmarking, Metrics, and Toolbox
di: Wang, Haohui, et al.
Pubblicazione: (2023)