Information-Theoretic Foundations for Neural Scaling Laws
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Jeon, Hong Jun, Van Roy, Benjamin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Information-Theoretic Foundations for Machine Learning
par: Jeon, Hong Jun, et autres
Publié: (2024)
par: Jeon, Hong Jun, et autres
Publié: (2024)
Aligning AI Agents via Information-Directed Sampling
par: Jeon, Hong Jun, et autres
Publié: (2024)
par: Jeon, Hong Jun, et autres
Publié: (2024)
The Need for a Big World Simulator: A Scientific Challenge for Continual Learning
par: Kumar, Saurabh, et autres
Publié: (2024)
par: Kumar, Saurabh, et autres
Publié: (2024)
Theoretical Foundations of Scaling Law in Familial Models
par: Song, Huan, et autres
Publié: (2025)
par: Song, Huan, et autres
Publié: (2025)
Towards Neural Scaling Laws for Time Series Foundation Models
par: Yao, Qingren, et autres
Publié: (2024)
par: Yao, Qingren, et autres
Publié: (2024)
Continual Learning as Computationally Constrained Reinforcement Learning
par: Kumar, Saurabh, et autres
Publié: (2023)
par: Kumar, Saurabh, et autres
Publié: (2023)
Scaling Laws and In-Context Learning: A Unified Theoretical Framework
par: Mehta, Sushant, et autres
Publié: (2025)
par: Mehta, Sushant, et autres
Publié: (2025)
Towards Neural Scaling Laws on Graphs
par: Liu, Jingzhe, et autres
Publié: (2024)
par: Liu, Jingzhe, et autres
Publié: (2024)
On the Optimizer Dependence of Neural Scaling Laws
par: Ramani, Vansh, et autres
Publié: (2026)
par: Ramani, Vansh, et autres
Publié: (2026)
An Information Theoretic Evaluation Metric For Strong Unlearning
par: Jeon, Dongjae, et autres
Publié: (2024)
par: Jeon, Dongjae, et autres
Publié: (2024)
Exploring Scaling Laws for EHR Foundation Models
par: Zhang, Sheng, et autres
Publié: (2025)
par: Zhang, Sheng, et autres
Publié: (2025)
Deriving Neural Scaling Laws from the statistics of natural language
par: Cagnetta, Francesco, et autres
Publié: (2026)
par: Cagnetta, Francesco, et autres
Publié: (2026)
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
par: Marklund, Henrik, et autres
Publié: (2024)
par: Marklund, Henrik, et autres
Publié: (2024)
Optimal Scaling Laws for Efficiency Gains in a Theoretical Transformer-Augmented Sectional MoE Framework
par: Sane, Soham
Publié: (2025)
par: Sane, Soham
Publié: (2025)
An Information-Theoretic Analysis of In-Context Learning
par: Jeon, Hong Jun, et autres
Publié: (2024)
par: Jeon, Hong Jun, et autres
Publié: (2024)
Do Neural Scaling Laws Exist on Graph Self-Supervised Learning?
par: Ma, Qian, et autres
Publié: (2024)
par: Ma, Qian, et autres
Publié: (2024)
Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks
par: Lee, Dongwoo, et autres
Publié: (2025)
par: Lee, Dongwoo, et autres
Publié: (2025)
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
par: Kothapalli, Vignesh, et autres
Publié: (2026)
par: Kothapalli, Vignesh, et autres
Publié: (2026)
A Theoretical Survey on Foundation Models
par: Fu, Shi, et autres
Publié: (2024)
par: Fu, Shi, et autres
Publié: (2024)
Efficient Exploration at Scale
par: Asghari, Seyed Mohammad, et autres
Publié: (2026)
par: Asghari, Seyed Mohammad, et autres
Publié: (2026)
The Scaling Law for LoRA Base on Mutual Information Upper Bound
par: Zhang, Jing, et autres
Publié: (2025)
par: Zhang, Jing, et autres
Publié: (2025)
Relative-Based Scaling Law for Neural Language Models
par: Yue, Baoqing, et autres
Publié: (2025)
par: Yue, Baoqing, et autres
Publié: (2025)
Effective Frontiers: A Unification of Neural Scaling Laws
par: Zou, Jiaxuan, et autres
Publié: (2026)
par: Zou, Jiaxuan, et autres
Publié: (2026)
Exploration Unbound
par: Arumugam, Dilip, et autres
Publié: (2024)
par: Arumugam, Dilip, et autres
Publié: (2024)
Consequentialist Objectives and Catastrophe
par: Marklund, Henrik, et autres
Publié: (2026)
par: Marklund, Henrik, et autres
Publié: (2026)
Maintaining Plasticity in Continual Learning via Regenerative Regularization
par: Kumar, Saurabh, et autres
Publié: (2023)
par: Kumar, Saurabh, et autres
Publié: (2023)
Misalignment from Treating Means as Ends
par: Marklund, Henrik, et autres
Publié: (2025)
par: Marklund, Henrik, et autres
Publié: (2025)
Adaptive Normalization Mamba with Multi Scale Trend Decomposition and Patch MoE Encoding
par: Jeon, MinCheol
Publié: (2025)
par: Jeon, MinCheol
Publié: (2025)
A Resource Model For Neural Scaling Law
par: Song, Jinyeop, et autres
Publié: (2024)
par: Song, Jinyeop, et autres
Publié: (2024)
The Neural Pruning Law Hypothesis
par: Barbulescu, Eugen, et autres
Publié: (2025)
par: Barbulescu, Eugen, et autres
Publié: (2025)
Scaling Laws of Motion Forecasting and Planning -- Technical Report
par: Baniodeh, Mustafa, et autres
Publié: (2025)
par: Baniodeh, Mustafa, et autres
Publié: (2025)
SCNO: Spiking Compositional Neural Operator -- Towards a Neuromorphic Foundation Model for Nuclear PDE Solving
par: Roy, Samrendra, et autres
Publié: (2026)
par: Roy, Samrendra, et autres
Publié: (2026)
Scaling Laws and Symmetry, Evidence from Neural Force Fields
par: Ngo, Khang, et autres
Publié: (2025)
par: Ngo, Khang, et autres
Publié: (2025)
On Implications of Scaling Laws on Feature Superposition
par: Katta, Pavan
Publié: (2024)
par: Katta, Pavan
Publié: (2024)
Scaling Law Hypothesis for Multimodal Model
par: Sun, Qingyun, et autres
Publié: (2024)
par: Sun, Qingyun, et autres
Publié: (2024)
Scaling Law for Time Series Forecasting
par: Shi, Jingzhe, et autres
Publié: (2024)
par: Shi, Jingzhe, et autres
Publié: (2024)
Wukong: Towards a Scaling Law for Large-Scale Recommendation
par: Zhang, Buyun, et autres
Publié: (2024)
par: Zhang, Buyun, et autres
Publié: (2024)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
par: Chen, Xiaodong, et autres
Publié: (2024)
par: Chen, Xiaodong, et autres
Publié: (2024)
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
par: Wu, Xiaojun, et autres
Publié: (2025)
par: Wu, Xiaojun, et autres
Publié: (2025)
Training Foundation Models as Data Compression: On Information, Model Weights and Copyright Law
par: Franceschelli, Giorgio, et autres
Publié: (2024)
par: Franceschelli, Giorgio, et autres
Publié: (2024)
Documents similaires
-
Information-Theoretic Foundations for Machine Learning
par: Jeon, Hong Jun, et autres
Publié: (2024) -
Aligning AI Agents via Information-Directed Sampling
par: Jeon, Hong Jun, et autres
Publié: (2024) -
The Need for a Big World Simulator: A Scientific Challenge for Continual Learning
par: Kumar, Saurabh, et autres
Publié: (2024) -
Theoretical Foundations of Scaling Law in Familial Models
par: Song, Huan, et autres
Publié: (2025) -
Towards Neural Scaling Laws for Time Series Foundation Models
par: Yao, Qingren, et autres
Publié: (2024)