Neural Neural Scaling Laws
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Michael Y., Pan, Jane, Jhaveri, Ayush Rajesh, Lourie, Nicholas, Cho, Kyunghyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
Show Your Work with Confidence: Confidence Bands for Tuning Curves
di: Lourie, Nicholas, et al.
Pubblicazione: (2023)
di: Lourie, Nicholas, et al.
Pubblicazione: (2023)
Hyperparameter Loss Surfaces Are Simple Near their Optima
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
di: Chen, Mayee F., et al.
Pubblicazione: (2024)
di: Chen, Mayee F., et al.
Pubblicazione: (2024)
Interpreting and Mitigating Unwanted Uncertainty in LLMs
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
di: Jhaveri, Ayush Rajesh, et al.
Pubblicazione: (2026)
di: Jhaveri, Ayush Rajesh, et al.
Pubblicazione: (2026)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
On the Relationship Between the Choice of Representation and In-Context Learning
di: Marinescu, Ioana, et al.
Pubblicazione: (2025)
di: Marinescu, Ioana, et al.
Pubblicazione: (2025)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
di: Park, Ji Won, et al.
Pubblicazione: (2025)
di: Park, Ji Won, et al.
Pubblicazione: (2025)
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
Relative-Based Scaling Law for Neural Language Models
di: Yue, Baoqing, et al.
Pubblicazione: (2025)
di: Yue, Baoqing, et al.
Pubblicazione: (2025)
Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable Modeling
di: Madaan, Divyam, et al.
Pubblicazione: (2026)
di: Madaan, Divyam, et al.
Pubblicazione: (2026)
Temporal Generalization: A Reality Check
di: Madaan, Divyam, et al.
Pubblicazione: (2025)
di: Madaan, Divyam, et al.
Pubblicazione: (2025)
Preference Learning Algorithms Do Not Learn Preference Rankings
di: Chen, Angelica, et al.
Pubblicazione: (2024)
di: Chen, Angelica, et al.
Pubblicazione: (2024)
Language Models as Causal Effect Generators
di: Bynum, Lucius E. J., et al.
Pubblicazione: (2024)
di: Bynum, Lucius E. J., et al.
Pubblicazione: (2024)
Training Language Models with Language Feedback at Scale
di: Scheurer, Jérémy, et al.
Pubblicazione: (2023)
di: Scheurer, Jérémy, et al.
Pubblicazione: (2023)
Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning
di: Madaan, Divyam, et al.
Pubblicazione: (2024)
di: Madaan, Divyam, et al.
Pubblicazione: (2024)
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
di: Madaan, Divyam, et al.
Pubblicazione: (2025)
di: Madaan, Divyam, et al.
Pubblicazione: (2025)
Scaling Laws for Precision
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
Scaling Laws For Mixed Quantization
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
di: Singh, Karan, et al.
Pubblicazione: (2026)
di: Singh, Karan, et al.
Pubblicazione: (2026)
MedLM: Exploring Language Models for Medical Question Answering Systems
di: Yagnik, Niraj, et al.
Pubblicazione: (2024)
di: Yagnik, Niraj, et al.
Pubblicazione: (2024)
CarbonScaling: Extending Neural Scaling Laws for Carbon Footprint in Large Language Models
di: Jiang, Lei, et al.
Pubblicazione: (2025)
di: Jiang, Lei, et al.
Pubblicazione: (2025)
Non-convolutional Graph Neural Networks
di: Wang, Yuanqing, et al.
Pubblicazione: (2024)
di: Wang, Yuanqing, et al.
Pubblicazione: (2024)
Unified Scaling Laws for Compressed Representations
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
Parallel Scaling Law for Language Models
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
Scaling Law for Quantization-Aware Training
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
Reconciling Kaplan and Chinchilla Scaling Laws
di: Pearce, Tim, et al.
Pubblicazione: (2024)
di: Pearce, Tim, et al.
Pubblicazione: (2024)
Scaling Laws for Multilingual Language Models
di: He, Yifei, et al.
Pubblicazione: (2024)
di: He, Yifei, et al.
Pubblicazione: (2024)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
Prescriptive Scaling Laws for Data Constrained Training
di: Lovelace, Justin, et al.
Pubblicazione: (2026)
di: Lovelace, Justin, et al.
Pubblicazione: (2026)
gzip Predicts Data-dependent Scaling Laws
di: Pandey, Rohan
Pubblicazione: (2024)
di: Pandey, Rohan
Pubblicazione: (2024)
Kinetics: Rethinking Test-Time Scaling Laws
di: Sadhukhan, Ranajoy, et al.
Pubblicazione: (2025)
di: Sadhukhan, Ranajoy, et al.
Pubblicazione: (2025)
Unraveling the Mystery of Scaling Laws: Part I
di: Su, Hui, et al.
Pubblicazione: (2024)
di: Su, Hui, et al.
Pubblicazione: (2024)
Compression Scaling Laws:Unifying Sparsity and Quantization
di: Frantar, Elias, et al.
Pubblicazione: (2025)
di: Frantar, Elias, et al.
Pubblicazione: (2025)
Superposition Yields Robust Neural Scaling
di: Liu, Yizhou, et al.
Pubblicazione: (2025)
di: Liu, Yizhou, et al.
Pubblicazione: (2025)
A Survey : Neural Networks for AMR-to-Text
di: Hao, Hongyu, et al.
Pubblicazione: (2022)
di: Hao, Hongyu, et al.
Pubblicazione: (2022)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
di: Chen, Yanxi, et al.
Pubblicazione: (2024)
di: Chen, Yanxi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
di: Lourie, Nicholas, et al.
Pubblicazione: (2025) -
Show Your Work with Confidence: Confidence Bands for Tuning Curves
di: Lourie, Nicholas, et al.
Pubblicazione: (2023) -
Hyperparameter Loss Surfaces Are Simple Near their Optima
di: Lourie, Nicholas, et al.
Pubblicazione: (2025) -
Aioli: A Unified Optimization Framework for Language Model Data Mixing
di: Chen, Mayee F., et al.
Pubblicazione: (2024) -
Interpreting and Mitigating Unwanted Uncertainty in LLMs
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)