Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
Fuente:
arXiv
Saved in:
| Main Authors: | Jha, Nandan Kumar, Reagen, Brandon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spectral Scaling Laws in Language Models: How Effectively Do Feed-Forward Networks Use Their Latent Space?
by: Jha, Nandan Kumar, et al.
Published: (2025)
by: Jha, Nandan Kumar, et al.
Published: (2025)
NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks
by: Jha, Nandan Kumar, et al.
Published: (2026)
by: Jha, Nandan Kumar, et al.
Published: (2026)
A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention
by: Jha, Nandan Kumar, et al.
Published: (2025)
by: Jha, Nandan Kumar, et al.
Published: (2025)
AERO: Entropy-Guided Framework for Private LLM Inference
by: Jha, Nandan Kumar, et al.
Published: (2024)
by: Jha, Nandan Kumar, et al.
Published: (2024)
ReLU's Revival: On the Entropic Overload in Normalization-Free Large Language Models
by: Jha, Nandan Kumar, et al.
Published: (2024)
by: Jha, Nandan Kumar, et al.
Published: (2024)
Entropy-Guided Attention for Private LLMs
by: Jha, Nandan Kumar, et al.
Published: (2025)
by: Jha, Nandan Kumar, et al.
Published: (2025)
DeepReShape: Redesigning Neural Networks for Efficient Private Inference
by: Jha, Nandan Kumar, et al.
Published: (2023)
by: Jha, Nandan Kumar, et al.
Published: (2023)
TruncFormer: Private LLM Inference Using Only Truncations
by: Yubeaton, Patrick, et al.
Published: (2024)
by: Yubeaton, Patrick, et al.
Published: (2024)
Network and Compiler Optimizations for Efficient Linear Algebra Kernels in Private Transformer Inference
by: Garimella, Karthik, et al.
Published: (2025)
by: Garimella, Karthik, et al.
Published: (2025)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
by: Kim, Juno, et al.
Published: (2026)
by: Kim, Juno, et al.
Published: (2026)
Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
by: Wan, Weilin, et al.
Published: (2026)
by: Wan, Weilin, et al.
Published: (2026)
Same Error, Different Function: The Optimizer as an Implicit Prior in Financial Time Series
by: Cortesi, Federico Vittorio, et al.
Published: (2026)
by: Cortesi, Federico Vittorio, et al.
Published: (2026)
Renormalizable Spectral-Shell Dynamics as the Origin of Neural Scaling Laws
by: Zhang, Yizhou
Published: (2025)
by: Zhang, Yizhou
Published: (2025)
Capacity-Aware Mixture Law Enables Efficient LLM Data Optimization
by: Li, Jingwei, et al.
Published: (2026)
by: Li, Jingwei, et al.
Published: (2026)
Towards Robust Scaling Laws for Optimizers
by: Volkova, Alexandra, et al.
Published: (2026)
by: Volkova, Alexandra, et al.
Published: (2026)
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
On the Optimizer Dependence of Neural Scaling Laws
by: Ramani, Vansh, et al.
Published: (2026)
by: Ramani, Vansh, et al.
Published: (2026)
Throughput Optimization as a Strategic Lever in Large-Scale AI Systems: Evidence from Dataloader and Memory Profiling Innovations
by: Jha, Mayank
Published: (2026)
by: Jha, Mayank
Published: (2026)
Same Target, Different Basins: Hard vs. Soft Labels for Annotator Distributions
by: Gheibi, Mirerfan, et al.
Published: (2026)
by: Gheibi, Mirerfan, et al.
Published: (2026)
Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
by: Shulgin, Egor, et al.
Published: (2026)
by: Shulgin, Egor, et al.
Published: (2026)
Complexity Scaling Laws for Neural Models using Combinatorial Optimization
by: Weissman, Lowell, et al.
Published: (2025)
by: Weissman, Lowell, et al.
Published: (2025)
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
by: Jiang, Yiding, et al.
Published: (2024)
by: Jiang, Yiding, et al.
Published: (2024)
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
by: Bian, Song, et al.
Published: (2025)
by: Bian, Song, et al.
Published: (2025)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
by: Ouyang, Xu, et al.
Published: (2026)
by: Ouyang, Xu, et al.
Published: (2026)
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
by: Allen-Zhu, Zeyuan, et al.
Published: (2024)
by: Allen-Zhu, Zeyuan, et al.
Published: (2024)
Scaling Laws for Precision
by: Kumar, Tanishq, et al.
Published: (2024)
by: Kumar, Tanishq, et al.
Published: (2024)
Quantum Circuit Simulation of Compartmental Drug Dynamics: Leveraging Variational Algorithms for Nonlinear Mixed-Effects Population Pharmacokinetics
by: Singh, Isshaan, et al.
Published: (2026)
by: Singh, Isshaan, et al.
Published: (2026)
Rank Is Not Capacity: Spectral Occupancy for Latent Graph Models
by: Nakis, Nikolaos, et al.
Published: (2026)
by: Nakis, Nikolaos, et al.
Published: (2026)
Same Graph, Different Likelihoods: Calibration of Autoregressive Graph Generators via Permutation-Equivalent Encodings
by: Fredsgaard, Laurits, et al.
Published: (2026)
by: Fredsgaard, Laurits, et al.
Published: (2026)
Scaling Laws are Redundancy Laws
by: Bi, Yuda, et al.
Published: (2025)
by: Bi, Yuda, et al.
Published: (2025)
Grokking vs. Learning: Same Features, Different Encodings
by: Manning-Coe, Dmitry, et al.
Published: (2025)
by: Manning-Coe, Dmitry, et al.
Published: (2025)
Improved Molecular Generation through Attribute-Driven Integrative Embeddings and GAN Selectivity
by: Joshi, Nandan, et al.
Published: (2025)
by: Joshi, Nandan, et al.
Published: (2025)
Different Paths, Same Destination: Designing New Physics-Inspired Dynamical Systems with Engineered Stability to Minimize the Ising Hamiltonian
by: Ekanayake, E. M. H. E. B., et al.
Published: (2025)
by: Ekanayake, E. M. H. E. B., et al.
Published: (2025)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
Two Facets of the Same Optimization Coin: Model Degradation and Representation Collapse in Graph Foundation Models
by: Li, Xunkai, et al.
Published: (2025)
by: Li, Xunkai, et al.
Published: (2025)
A Comprehensively Adaptive Architectural Optimization-Ingrained Quantum Neural Network Model for Cloud Workloads Prediction
by: Kumar, Jitendra, et al.
Published: (2025)
by: Kumar, Jitendra, et al.
Published: (2025)
On the Invariance and Generality of Neural Scaling Laws
by: Han, Xing, et al.
Published: (2026)
by: Han, Xing, et al.
Published: (2026)
Scaling Laws of Global Weather Models
by: Yu, Yuejiang, et al.
Published: (2026)
by: Yu, Yuejiang, et al.
Published: (2026)
Towards Scaling Laws for Symbolic Regression
by: Otte, David, et al.
Published: (2025)
by: Otte, David, et al.
Published: (2025)
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
Similar Items
-
Spectral Scaling Laws in Language Models: How Effectively Do Feed-Forward Networks Use Their Latent Space?
by: Jha, Nandan Kumar, et al.
Published: (2025) -
NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks
by: Jha, Nandan Kumar, et al.
Published: (2026) -
A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention
by: Jha, Nandan Kumar, et al.
Published: (2025) -
AERO: Entropy-Guided Framework for Private LLM Inference
by: Jha, Nandan Kumar, et al.
Published: (2024) -
ReLU's Revival: On the Entropic Overload in Normalization-Free Large Language Models
by: Jha, Nandan Kumar, et al.
Published: (2024)