Small-to-Large Generalization: Data Influences Models Consistently Across Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Khaddaj, Alaa, Engstrom, Logan, Madry, Aleksander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DataMIL: Selecting Data for Robot Imitation Learning with Datamodels
by: Dass, Shivin, et al.
Published: (2025)
by: Dass, Shivin, et al.
Published: (2025)
DsDm: Model-Aware Dataset Selection with Datamodels
by: Engstrom, Logan, et al.
Published: (2024)
by: Engstrom, Logan, et al.
Published: (2024)
Optimizing ML Training with Metagradient Descent
by: Engstrom, Logan, et al.
Published: (2025)
by: Engstrom, Logan, et al.
Published: (2025)
Do Large Language Model Benchmarks Test Reliability?
by: Vendrow, Joshua, et al.
Published: (2025)
by: Vendrow, Joshua, et al.
Published: (2025)
Decomposing and Editing Predictions by Modeling Model Computation
by: Shah, Harshay, et al.
Published: (2024)
by: Shah, Harshay, et al.
Published: (2024)
ContextCite: Attributing Model Generation to Context
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
Large-Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season
by: Cen, Sarah H., et al.
Published: (2025)
by: Cen, Sarah H., et al.
Published: (2025)
Ask Your Distribution Shift if Pre-Training is Right for You
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
MAGIC: Near-Optimal Data Attribution for Deep Learning
by: Ilyas, Andrew, et al.
Published: (2025)
by: Ilyas, Andrew, et al.
Published: (2025)
Learning to Attribute with Attention
by: Cohen-Wang, Benjamin, et al.
Published: (2025)
by: Cohen-Wang, Benjamin, et al.
Published: (2025)
Data Debiasing with Datamodels (D3M): Improving Subgroup Robustness via Data Selection
by: Jain, Saachi, et al.
Published: (2024)
by: Jain, Saachi, et al.
Published: (2024)
User Strategization and Trustworthy Algorithms
by: Cen, Sarah H., et al.
Published: (2023)
by: Cen, Sarah H., et al.
Published: (2023)
Decomposition of Small Transformer Models
by: Christensen, Casper L., et al.
Published: (2025)
by: Christensen, Casper L., et al.
Published: (2025)
AI Supply Chains: An Emerging Ecosystem of AI Actors, Products, and Services
by: Hopkins, Aspen, et al.
Published: (2025)
by: Hopkins, Aspen, et al.
Published: (2025)
Attribute-to-Delete: Machine Unlearning via Datamodel Matching
by: Georgiev, Kristian, et al.
Published: (2024)
by: Georgiev, Kristian, et al.
Published: (2024)
LLM Circuit Analyses Are Consistent Across Training and Scale
by: Tigges, Curt, et al.
Published: (2024)
by: Tigges, Curt, et al.
Published: (2024)
Synthetic Data Generation for Augmenting Small Samples
by: Liu, Dan, et al.
Published: (2025)
by: Liu, Dan, et al.
Published: (2025)
Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
by: Qiu, Shikai, et al.
Published: (2025)
by: Qiu, Shikai, et al.
Published: (2025)
Measuring Strategization in Recommendation: Users Adapt Their Behavior to Shape Future Content
by: Cen, Sarah H., et al.
Published: (2024)
by: Cen, Sarah H., et al.
Published: (2024)
Veridical Data Science for Medical Foundation Models
by: Alaa, Ahmed, et al.
Published: (2024)
by: Alaa, Ahmed, et al.
Published: (2024)
Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models
by: Gilhuly, Colleen, et al.
Published: (2025)
by: Gilhuly, Colleen, et al.
Published: (2025)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
by: Zhang, Zishi, et al.
Published: (2026)
by: Zhang, Zishi, et al.
Published: (2026)
Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives?
by: Chakraborty, Sanjay, et al.
Published: (2025)
by: Chakraborty, Sanjay, et al.
Published: (2025)
Assessing the Potential of Masked Autoencoder Foundation Models in Predicting Downhole Metrics from Surface Drilling Data
by: Berezowski, Aleksander, et al.
Published: (2026)
by: Berezowski, Aleksander, et al.
Published: (2026)
Convergence Of Consistency Model With Multistep Sampling Under General Data Assumptions
by: Chen, Yiding, et al.
Published: (2025)
by: Chen, Yiding, et al.
Published: (2025)
Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching
by: Khamis, Alaa, et al.
Published: (2026)
by: Khamis, Alaa, et al.
Published: (2026)
Dimensionality Reduction on Riemannian Manifolds in Data Analysis
by: Ichi, Alaa El, et al.
Published: (2026)
by: Ichi, Alaa El, et al.
Published: (2026)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
by: Majumdar, Somshubra, et al.
Published: (2024)
by: Majumdar, Somshubra, et al.
Published: (2024)
Can Small-Scale Data Poisoning Exacerbate Dialect-Linked Biases in Large Language Models?
by: Abbas, Chaymaa, et al.
Published: (2025)
by: Abbas, Chaymaa, et al.
Published: (2025)
Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models
by: Lu, Cheng, et al.
Published: (2024)
by: Lu, Cheng, et al.
Published: (2024)
Fast Partition-Based Cross-Validation With Centering and Scaling for $\mathbf{X}^\mathbf{T}\mathbf{X}$ and $\mathbf{X}^\mathbf{T}\mathbf{Y}$
by: Engstrøm, Ole-Christian Galbo, et al.
Published: (2024)
by: Engstrøm, Ole-Christian Galbo, et al.
Published: (2024)
Generalizing Large Language Model Usability Across Resource-Constrained
by: Tsai, Yun-Da
Published: (2025)
by: Tsai, Yun-Da
Published: (2025)
Can Large Language Models Generalize Procedures Across Representations?
by: Lin, Fangru, et al.
Published: (2026)
by: Lin, Fangru, et al.
Published: (2026)
DAD4TS: Data-Augmentation-Oriented Diffusion Model for Time-Series Forecasting with Small-Scale Data
by: Suzuki, Masahiro, et al.
Published: (2026)
by: Suzuki, Masahiro, et al.
Published: (2026)
NeuralSolver: Learning Algorithms For Consistent and Efficient Extrapolation Across General Tasks
by: Esteves, Bernardo, et al.
Published: (2024)
by: Esteves, Bernardo, et al.
Published: (2024)
Data Selection: A General Principle for Building Small Interpretable Models
by: Ghose, Abhishek
Published: (2022)
by: Ghose, Abhishek
Published: (2022)
Area Modeling using Stay Information for Large-Scale Users and Analysis for Influence of COVID-19
by: Shoji, Kazuyuki, et al.
Published: (2024)
by: Shoji, Kazuyuki, et al.
Published: (2024)
Efficient Data Subset Selection to Generalize Training Across Models: Transductive and Inductive Networks
by: Jain, Eeshaan, et al.
Published: (2024)
by: Jain, Eeshaan, et al.
Published: (2024)
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
by: Bogavelli, Tara, et al.
Published: (2026)
by: Bogavelli, Tara, et al.
Published: (2026)
Layer-wise Lipschitz-Product Control for Deep Kolmogorov--Arnold Network Representations of Compositionally Structured Functions
by: Tankman, Aleksander
Published: (2026)
by: Tankman, Aleksander
Published: (2026)
Similar Items
-
DataMIL: Selecting Data for Robot Imitation Learning with Datamodels
by: Dass, Shivin, et al.
Published: (2025) -
DsDm: Model-Aware Dataset Selection with Datamodels
by: Engstrom, Logan, et al.
Published: (2024) -
Optimizing ML Training with Metagradient Descent
by: Engstrom, Logan, et al.
Published: (2025) -
Do Large Language Model Benchmarks Test Reliability?
by: Vendrow, Joshua, et al.
Published: (2025) -
Decomposing and Editing Predictions by Modeling Model Computation
by: Shah, Harshay, et al.
Published: (2024)