Effective Frontiers: A Unification of Neural Scaling Laws
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Jiaxuan, Gong, Zixuan, Su, Ye, Tang, Huayi, Liu, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Capabilities and Fundamental Limits of Latent Chain-of-Thought
by: Zou, Jiaxuan, et al.
Published: (2026)
by: Zou, Jiaxuan, et al.
Published: (2026)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Generative AI and Process Systems Engineering: The Next Frontier
by: Decardi-Nelson, Benjamin, et al.
Published: (2024)
by: Decardi-Nelson, Benjamin, et al.
Published: (2024)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
Deep Learning for Sequential Decision Making under Uncertainty: Foundations, Frameworks, and Frontiers
by: Buyuktahtakin, I. Esra
Published: (2026)
by: Buyuktahtakin, I. Esra
Published: (2026)
Constructing Industrial-Scale Optimization Modeling Benchmark
by: Li, Zhong, et al.
Published: (2026)
by: Li, Zhong, et al.
Published: (2026)
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
by: Su, Weijie
Published: (2025)
by: Su, Weijie
Published: (2025)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
by: Choudhury, Sayantan, et al.
Published: (2024)
by: Choudhury, Sayantan, et al.
Published: (2024)
Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
by: Ghosh, Ipsita, et al.
Published: (2025)
by: Ghosh, Ipsita, et al.
Published: (2025)
ARO: A New Lens On Matrix Optimization For Large Models
by: Gong, Wenbo, et al.
Published: (2026)
by: Gong, Wenbo, et al.
Published: (2026)
Neural Solver Selection for Combinatorial Optimization
by: Gao, Chengrui, et al.
Published: (2024)
by: Gao, Chengrui, et al.
Published: (2024)
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
by: Hrycej, Tomas, et al.
Published: (2025)
by: Hrycej, Tomas, et al.
Published: (2025)
Optimal Control Operator Perspective and a Neural Adaptive Spectral Method
by: Feng, Mingquan, et al.
Published: (2024)
by: Feng, Mingquan, et al.
Published: (2024)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
Bayesian Optimization for Hyperparameters Tuning in Neural Networks
by: Onorato, Gabriele
Published: (2024)
by: Onorato, Gabriele
Published: (2024)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
by: Shen, Jucheng, et al.
Published: (2026)
by: Shen, Jucheng, et al.
Published: (2026)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
Training Safe Neural Networks with Global SDP Bounds
by: Soletskyi, Roman, et al.
Published: (2024)
by: Soletskyi, Roman, et al.
Published: (2024)
Neur2BiLO: Neural Bilevel Optimization
by: Dumouchelle, Justin, et al.
Published: (2024)
by: Dumouchelle, Justin, et al.
Published: (2024)
Applications of 0-1 Neural Networks in Prescription and Prediction
by: Patil, Vrishabh, et al.
Published: (2024)
by: Patil, Vrishabh, et al.
Published: (2024)
Taming Binarized Neural Networks and Mixed-Integer Programs
by: Aspman, Johannes, et al.
Published: (2023)
by: Aspman, Johannes, et al.
Published: (2023)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
by: Han, X. Y., et al.
Published: (2025)
by: Han, X. Y., et al.
Published: (2025)
Towards Efficient Constraint Handling in Neural Solvers for Routing Problems
by: Bi, Jieyi, et al.
Published: (2026)
by: Bi, Jieyi, et al.
Published: (2026)
High-order expansion of Neural Ordinary Differential Equations flows
by: Izzo, Dario, et al.
Published: (2025)
by: Izzo, Dario, et al.
Published: (2025)
Feed-Forward Neural Networks as a Mixed-Integer Program
by: Aftabi, Navid, et al.
Published: (2024)
by: Aftabi, Navid, et al.
Published: (2024)
Graph Neural Networks for the Offline Nanosatellite Task Scheduling Problem
by: Pacheco, Bruno Machado, et al.
Published: (2023)
by: Pacheco, Bruno Machado, et al.
Published: (2023)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
LLM Serving Optimization with Variable Prefill and Decode Lengths
by: Wang, Meixuan, et al.
Published: (2025)
by: Wang, Meixuan, et al.
Published: (2025)
Self-Certifying Primal-Dual Optimization Proxies for Large-Scale Batch Economic Dispatch
by: Klamkin, Michael, et al.
Published: (2025)
by: Klamkin, Michael, et al.
Published: (2025)
Is Scaling Learned Optimizers Worth It? Evaluating The Value of VeLO's 4000 TPU Months
by: Rezk, Fady, et al.
Published: (2023)
by: Rezk, Fady, et al.
Published: (2023)
SMiLE: Provably Enforcing Global Relational Properties in Neural Networks
by: Francobaldi, Matteo, et al.
Published: (2025)
by: Francobaldi, Matteo, et al.
Published: (2025)
Optimizing the Optimizer for Physics-Informed Neural Networks and Kolmogorov-Arnold Networks
by: Kiyani, Elham, et al.
Published: (2025)
by: Kiyani, Elham, et al.
Published: (2025)
Ginger: An Efficient Curvature Approximation with Linear Complexity for General Neural Networks
by: Hao, Yongchang, et al.
Published: (2024)
by: Hao, Yongchang, et al.
Published: (2024)
Neural Combinatorial Optimization for Stochastic Flexible Job Shop Scheduling Problems
by: Smit, Igor G., et al.
Published: (2024)
by: Smit, Igor G., et al.
Published: (2024)
Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial
by: Li, Haoyu, et al.
Published: (2026)
by: Li, Haoyu, et al.
Published: (2026)
A Riemannian Optimization Perspective of the Gauss-Newton Method for Feedforward Neural Networks
by: Cayci, Semih
Published: (2024)
by: Cayci, Semih
Published: (2024)
ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule
by: Huang, Yilie, et al.
Published: (2026)
by: Huang, Yilie, et al.
Published: (2026)
Unsupervised Training of Diffusion Models for Feasible Solution Generation in Neural Combinatorial Optimization
by: Hong, Seong-Hyun, et al.
Published: (2024)
by: Hong, Seong-Hyun, et al.
Published: (2024)
Similar Items
-
Capabilities and Fundamental Limits of Latent Chain-of-Thought
by: Zou, Jiaxuan, et al.
Published: (2026) -
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026) -
Generative AI and Process Systems Engineering: The Next Frontier
by: Decardi-Nelson, Benjamin, et al.
Published: (2024) -
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024) -
Deep Learning for Sequential Decision Making under Uncertainty: Foundations, Frameworks, and Frontiers
by: Buyuktahtakin, I. Esra
Published: (2026)