Why Less is More (Sometimes): A Theory of Data Curation
Fuente:
arXiv
Saved in:
| Main Authors: | Dohmatob, Elvis, Pezeshki, Mohammad, Askari-Hemmat, Reyhane |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
Feedback-guided Data Synthesis for Imbalanced Classification
by: Hemmat, Reyhane Askari, et al.
Published: (2023)
by: Hemmat, Reyhane Askari, et al.
Published: (2023)
auto-fpt: Automating Free Probability Theory Calculations for Machine Learning Theory
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
The Pitfalls of Memorization: When Memorization Hurts Generalization
by: Bayat, Reza, et al.
Published: (2024)
by: Bayat, Reza, et al.
Published: (2024)
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
by: Pasand, Ali Saheb, et al.
Published: (2025)
by: Pasand, Ali Saheb, et al.
Published: (2025)
QGen: On the Ability to Generalize in Quantization Aware Training
by: AskariHemmat, MohammadHossein, et al.
Published: (2024)
by: AskariHemmat, MohammadHossein, et al.
Published: (2024)
(Sometimes) Less is More: Mitigating the Complexity of Rule-based Representation for Interpretable Classification
by: Bergamin, Luca, et al.
Published: (2025)
by: Bergamin, Luca, et al.
Published: (2025)
An Effective Theory of Bias Amplification
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
Model Collapse Demystified: The Case of Regression
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Efficient Refusal Ablation in LLM through Optimal Transport
by: Nanfack, Geraldin, et al.
Published: (2026)
by: Nanfack, Geraldin, et al.
Published: (2026)
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
by: Hemmat, Reyhane Askari, et al.
Published: (2024)
by: Hemmat, Reyhane Askari, et al.
Published: (2024)
Strong Model Collapse
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification
by: Feng, Yunzhen, et al.
Published: (2024)
by: Feng, Yunzhen, et al.
Published: (2024)
Scaling Laws for Associative Memories
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Less is More: Adaptive Coverage for Synthetic Training Data
by: Tavakkol, Sasan, et al.
Published: (2025)
by: Tavakkol, Sasan, et al.
Published: (2025)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Why Zeroth-Order Adaptation May Forget Less: A Randomized Shaping Theory
by: Shu, Yao, et al.
Published: (2026)
by: Shu, Yao, et al.
Published: (2026)
Steinmetz Neural Networks for Complex-Valued Data
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
by: Gong, Shuzhi, et al.
Published: (2026)
by: Gong, Shuzhi, et al.
Published: (2026)
Delta-Audit: Explaining What Changes When Models Change
by: Hemmat, Arshia, et al.
Published: (2025)
by: Hemmat, Arshia, et al.
Published: (2025)
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
by: Öncel, Fırat, et al.
Published: (2024)
by: Öncel, Fırat, et al.
Published: (2024)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
The Alignment Game: A Theory of Long-Horizon Alignment Through Recursive Curation
by: Falahati, Ali, et al.
Published: (2025)
by: Falahati, Ali, et al.
Published: (2025)
Transformer Multivariate Forecasting: Less is More?
by: Xu, Jingjing, et al.
Published: (2023)
by: Xu, Jingjing, et al.
Published: (2023)
Continual Learning: Less Forgetting, More OOD Generalization via Adaptive Contrastive Replay
by: Rezaei, Hossein, et al.
Published: (2024)
by: Rezaei, Hossein, et al.
Published: (2024)
Quantize What Counts: More for Keys, Less for Values
by: Hariri, Mohsen, et al.
Published: (2025)
by: Hariri, Mohsen, et al.
Published: (2025)
Less is More: Towards Simple Graph Contrastive Learning
by: Zhao, Yanan, et al.
Published: (2025)
by: Zhao, Yanan, et al.
Published: (2025)
Self-Ablating Transformers: More Interpretability, Less Sparsity
by: Ferrao, Jeremias, et al.
Published: (2025)
by: Ferrao, Jeremias, et al.
Published: (2025)
Reasoning Models Sometimes Output Illegible Chains of Thought
by: Jose, Arun
Published: (2025)
by: Jose, Arun
Published: (2025)
Iterative Amortized Inference: Unifying In-Context Learning and Learned Optimizers
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Leveraging Retrieval-Augmented Generation for Persian University Knowledge Retrieval
by: Hemmat, Arshia, et al.
Published: (2024)
by: Hemmat, Arshia, et al.
Published: (2024)
Context Awareness Gate For Retrieval Augmented Generation
by: Heydari, Mohammad Hassan, et al.
Published: (2024)
by: Heydari, Mohammad Hassan, et al.
Published: (2024)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
by: Ye, Jiayuan, et al.
Published: (2026)
by: Ye, Jiayuan, et al.
Published: (2026)
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
by: Qin, Tian, et al.
Published: (2024)
by: Qin, Tian, et al.
Published: (2024)
Achieving More with Less: A Tensor-Optimization-Powered Ensemble Method
by: Yuan, Jinghui, et al.
Published: (2024)
by: Yuan, Jinghui, et al.
Published: (2024)
Less is More: Recursive Reasoning with Tiny Networks
by: Jolicoeur-Martineau, Alexia
Published: (2025)
by: Jolicoeur-Martineau, Alexia
Published: (2025)
No More, No Less: Least-Privilege Language Models
by: Rauba, Paulius, et al.
Published: (2026)
by: Rauba, Paulius, et al.
Published: (2026)
RL's Razor: Why Online Reinforcement Learning Forgets Less
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
Rethinking Tokenization for Clinical Time Series: When Less is More
by: Attrach, Rafi Al, et al.
Published: (2025)
by: Attrach, Rafi Al, et al.
Published: (2025)
Similar Items
-
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
by: Askari-Hemmat, Reyhane, et al.
Published: (2025) -
Feedback-guided Data Synthesis for Imbalanced Classification
by: Hemmat, Reyhane Askari, et al.
Published: (2023) -
auto-fpt: Automating Free Probability Theory Calculations for Machine Learning Theory
by: Subramonian, Arjun, et al.
Published: (2025) -
The Pitfalls of Memorization: When Memorization Hurts Generalization
by: Bayat, Reza, et al.
Published: (2024) -
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
by: Pasand, Ali Saheb, et al.
Published: (2025)