From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Kothapalli, Vignesh, Pang, Tianyu, Deng, Shenyang, Liu, Zongmin, Yang, Yaoqing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
Can Kernel Methods Explain How the Data Affects Neural Collapse?
by: Kothapalli, Vignesh, et al.
Published: (2024)
by: Kothapalli, Vignesh, et al.
Published: (2024)
Chemical Reaction Networks Learn Better than Spiking Neural Networks
by: Jaffard, Sophie, et al.
Published: (2026)
by: Jaffard, Sophie, et al.
Published: (2026)
Neural Networks Generalize on Low Complexity Data
by: Chatterjee, Sourav, et al.
Published: (2024)
by: Chatterjee, Sourav, et al.
Published: (2024)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
by: Zhang, Huiming, et al.
Published: (2026)
by: Zhang, Huiming, et al.
Published: (2026)
Phase-Type Variational Autoencoders for Heavy-Tailed Data
by: Ziani, Abdelhakim, et al.
Published: (2026)
by: Ziani, Abdelhakim, et al.
Published: (2026)
Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive Ability
by: Yu, Lijia, et al.
Published: (2025)
by: Yu, Lijia, et al.
Published: (2025)
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon
by: Yu, Hao
Published: (2025)
by: Yu, Hao
Published: (2025)
Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
by: Hu, Xinyang, et al.
Published: (2024)
by: Hu, Xinyang, et al.
Published: (2024)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
Convergence of Shallow ReLU Networks on Weakly Interacting Data
by: Dana, Léo, et al.
Published: (2025)
by: Dana, Léo, et al.
Published: (2025)
Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks
by: Fu, Hengyu, et al.
Published: (2024)
by: Fu, Hengyu, et al.
Published: (2024)
A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
by: Yu, Hao
Published: (2026)
by: Yu, Hao
Published: (2026)
Asymptotic Classification Error for Heavy-Tailed Renewal Processes
by: Rong, Xinhui, et al.
Published: (2024)
by: Rong, Xinhui, et al.
Published: (2024)
Spectrally-Corrected and Regularized QDA Classifier for Spiked Covariance Model
by: Luo, Wenya, et al.
Published: (2025)
by: Luo, Wenya, et al.
Published: (2025)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
by: Su, Xun, et al.
Published: (2025)
by: Su, Xun, et al.
Published: (2025)
How Particle-System Random Batch Methods Enhance Graph Transformer: Memory Efficiency and Parallel Computing Strategy
by: Liu, Hanwen, et al.
Published: (2025)
by: Liu, Hanwen, et al.
Published: (2025)
Provable Robust Overfitting Mitigation in Wasserstein Distributionally Robust Optimization
by: Liu, Shuang, et al.
Published: (2025)
by: Liu, Shuang, et al.
Published: (2025)
Conformal Policy Control
by: Prinster, Drew, et al.
Published: (2026)
by: Prinster, Drew, et al.
Published: (2026)
Sparsified-Learning for High-Dimensional Heavy-Tailed Locally Stationary Time Series, Concentration and Oracle Inequalities
by: Wang, Yingjie, et al.
Published: (2025)
by: Wang, Yingjie, et al.
Published: (2025)
CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations
by: Kothapalli, Vignesh, et al.
Published: (2025)
by: Kothapalli, Vignesh, et al.
Published: (2025)
Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation
by: Liu, Xiaotian, et al.
Published: (2026)
by: Liu, Xiaotian, et al.
Published: (2026)
A Likelihood Based Approach to Distribution Regression Using Conditional Deep Generative Models
by: Kumar, Shivam, et al.
Published: (2024)
by: Kumar, Shivam, et al.
Published: (2024)
A Separation in Heavy-Tailed Sampling: Gaussian vs. Stable Oracles for Proximal Samplers
by: He, Ye, et al.
Published: (2024)
by: He, Ye, et al.
Published: (2024)
Optimal rates for density and mode estimation with expand-and-sparsify representations
by: Sinha, Kaushik, et al.
Published: (2026)
by: Sinha, Kaushik, et al.
Published: (2026)
Influence functions and regularity tangents for efficient active learning
by: Eaton, Frederik
Published: (2024)
by: Eaton, Frederik
Published: (2024)
Efficient Knowledge Distillation via Curriculum Extraction
by: Gupta, Shivam, et al.
Published: (2025)
by: Gupta, Shivam, et al.
Published: (2025)
Risk Analysis and Design Against Adversarial Actions
by: Campi, Marco C., et al.
Published: (2025)
by: Campi, Marco C., et al.
Published: (2025)
A Theory of the Mechanics of Information: Generalization Through Measurement of Uncertainty (Learning is Measuring)
by: Hazard, Christopher J., et al.
Published: (2025)
by: Hazard, Christopher J., et al.
Published: (2025)
Solving a Research Problem in Mathematical Statistics with AI Assistance
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
Generalising realisability in statistical learning theory under epistemic uncertainty
by: Cuzzolin, Fabio
Published: (2024)
by: Cuzzolin, Fabio
Published: (2024)
Training Implicit Generative Models via an Invariant Statistical Loss
by: de Frutos, José Manuel, et al.
Published: (2024)
by: de Frutos, José Manuel, et al.
Published: (2024)
Identifiability of Potentially Degenerate Gaussian Mixture Models With Piecewise Affine Mixing
by: Xu, Danru, et al.
Published: (2026)
by: Xu, Danru, et al.
Published: (2026)
Cross-regularization: Adaptive Model Complexity through Validation Gradients
by: Brito, Carlos Stein
Published: (2025)
by: Brito, Carlos Stein
Published: (2025)
Residual Feature Integration is Sufficient to Prevent Negative Transfer
by: Xu, Yichen, et al.
Published: (2025)
by: Xu, Yichen, et al.
Published: (2025)
Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data
by: Chakraborty, Saptarshi, et al.
Published: (2026)
by: Chakraborty, Saptarshi, et al.
Published: (2026)
A Fine-Grained Understanding of Uniform Convergence for Halfspaces
by: Kontorovich, Aryeh, et al.
Published: (2026)
by: Kontorovich, Aryeh, et al.
Published: (2026)
Compression, Generalization and Learning
by: Campi, Marco C., et al.
Published: (2023)
by: Campi, Marco C., et al.
Published: (2023)
Similar Items
-
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
by: Pang, Tianyu, et al.
Published: (2026) -
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
by: Pang, Tianyu, et al.
Published: (2026) -
Can Kernel Methods Explain How the Data Affects Neural Collapse?
by: Kothapalli, Vignesh, et al.
Published: (2024) -
Chemical Reaction Networks Learn Better than Spiking Neural Networks
by: Jaffard, Sophie, et al.
Published: (2026) -
Neural Networks Generalize on Low Complexity Data
by: Chatterjee, Sourav, et al.
Published: (2024)