LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
Fuente:
arXiv
Salvato in:
| Autori principali: | Mayilvahanan, Prasanna, Wiedemer, Thaddäus, Mallick, Sayak, Bethge, Matthias, Brendel, Wieland |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2023)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2023)
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
di: Zeller, Jana, et al.
Pubblicazione: (2026)
di: Zeller, Jana, et al.
Pubblicazione: (2026)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025)
In Search of Forgotten Domain Generalization
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2024)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2024)
Provable Compositional Generalization for Object-Centric Learning
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2023)
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2023)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
VGGSounder: Audio-Visual Evaluations for Foundation Models
di: Zverev, Daniil, et al.
Pubblicazione: (2025)
di: Zverev, Daniil, et al.
Pubblicazione: (2025)
Mapping Post-Training Forgetting in Language Models at Scale
di: Harmon, Jackson, et al.
Pubblicazione: (2025)
di: Harmon, Jackson, et al.
Pubblicazione: (2025)
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
di: Öncel, Fırat, et al.
Pubblicazione: (2024)
di: Öncel, Fırat, et al.
Pubblicazione: (2024)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
di: Luo, Kairong, et al.
Pubblicazione: (2025)
di: Luo, Kairong, et al.
Pubblicazione: (2025)
Scaling Laws for Predicting Downstream Performance in LLMs
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
LLM Unlearning via Loss Adjustment with Only Forget Data
di: Wang, Yaxuan, et al.
Pubblicazione: (2024)
di: Wang, Yaxuan, et al.
Pubblicazione: (2024)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
di: Bergsma, Shane, et al.
Pubblicazione: (2025)
di: Bergsma, Shane, et al.
Pubblicazione: (2025)
LLM generation novelty through the lens of semantic similarity
di: Davydov, Philipp, et al.
Pubblicazione: (2025)
di: Davydov, Philipp, et al.
Pubblicazione: (2025)
On the Effect of Instruction Tuning Loss on Generalization
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2025)
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2025)
Loss Landscape Degeneracy and Stagewise Development in Transformers
di: Hoogland, Jesse, et al.
Pubblicazione: (2024)
di: Hoogland, Jesse, et al.
Pubblicazione: (2024)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
di: Sinha, Shiven, et al.
Pubblicazione: (2024)
di: Sinha, Shiven, et al.
Pubblicazione: (2024)
Towards Practical Tool Usage for Continually Learning LLMs
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs
di: Schuhmann, Christoph, et al.
Pubblicazione: (2025)
di: Schuhmann, Christoph, et al.
Pubblicazione: (2025)
Distillation Scaling Laws
di: Busbridge, Dan, et al.
Pubblicazione: (2025)
di: Busbridge, Dan, et al.
Pubblicazione: (2025)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
Instruction Fine-Tuning: Does Prompt Loss Matter?
di: Huerta-Enochian, Mathew, et al.
Pubblicazione: (2024)
di: Huerta-Enochian, Mathew, et al.
Pubblicazione: (2024)
What Scales in Cross-Entropy Scaling Law?
di: Yan, Junxi, et al.
Pubblicazione: (2025)
di: Yan, Junxi, et al.
Pubblicazione: (2025)
Understanding Emergent Abilities of Language Models from the Loss Perspective
di: Du, Zhengxiao, et al.
Pubblicazione: (2024)
di: Du, Zhengxiao, et al.
Pubblicazione: (2024)
G-Loss: Graph-Guided Fine-Tuning of Language Models
di: Sharma, Aditya, et al.
Pubblicazione: (2026)
di: Sharma, Aditya, et al.
Pubblicazione: (2026)
A Pseudo-Semantic Loss for Autoregressive Models with Logical Constraints
di: Ahmed, Kareem, et al.
Pubblicazione: (2023)
di: Ahmed, Kareem, et al.
Pubblicazione: (2023)
Fine-Tuning a Time Series Foundation Model with Wasserstein Loss
di: Chernov, Andrei
Pubblicazione: (2024)
di: Chernov, Andrei
Pubblicazione: (2024)
Scaling Law with Learning Rate Annealing
di: Tissue, Howe, et al.
Pubblicazione: (2024)
di: Tissue, Howe, et al.
Pubblicazione: (2024)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
di: Zeng, Liang, et al.
Pubblicazione: (2024)
di: Zeng, Liang, et al.
Pubblicazione: (2024)
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
di: Xie, Tong, et al.
Pubblicazione: (2025)
di: Xie, Tong, et al.
Pubblicazione: (2025)
Can Language Models Discover Scaling Laws?
di: Lin, Haowei, et al.
Pubblicazione: (2025)
di: Lin, Haowei, et al.
Pubblicazione: (2025)
Exploring Scaling Laws for EHR Foundation Models
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
Theoretical Foundations of Scaling Law in Familial Models
di: Song, Huan, et al.
Pubblicazione: (2025)
di: Song, Huan, et al.
Pubblicazione: (2025)
Scaling Laws for Fine-Grained Mixture of Experts
di: Krajewski, Jakub, et al.
Pubblicazione: (2024)
di: Krajewski, Jakub, et al.
Pubblicazione: (2024)
A Hitchhiker's Guide to Scaling Law Estimation
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
Relation-Aware Network with Attention-Based Loss for Few-Shot Knowledge Graph Completion
di: Qiao, Qiao, et al.
Pubblicazione: (2023)
di: Qiao, Qiao, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2023) -
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025) -
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
di: Zeller, Jana, et al.
Pubblicazione: (2026) -
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
di: Wiedemer, Thaddäus, et al.
Pubblicazione: (2025) -
In Search of Forgotten Domain Generalization
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2024)