Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Tianyu, Kothapalli, Vignesh, Deng, Shenyang, Wang, Haohui, Zhou, Dawei, Yang, Yaoqing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
by: Kothapalli, Vignesh, et al.
Published: (2024)
by: Kothapalli, Vignesh, et al.
Published: (2024)
Suspicious Alignment of SGD: A Fine-Grained Step Size Condition Analysis
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
KCES: Training-Free Defense for Robust Graph Neural Networks via Kernel Complexity
by: Jia, Yaning, et al.
Published: (2025)
by: Jia, Yaning, et al.
Published: (2025)
Can Kernel Methods Explain How the Data Affects Neural Collapse?
by: Kothapalli, Vignesh, et al.
Published: (2024)
by: Kothapalli, Vignesh, et al.
Published: (2024)
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
Model Balancing Helps Low-data Training and Fine-tuning
by: Liu, Zihang, et al.
Published: (2024)
by: Liu, Zihang, et al.
Published: (2024)
CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations
by: Kothapalli, Vignesh, et al.
Published: (2025)
by: Kothapalli, Vignesh, et al.
Published: (2025)
EvoluNet: Advancing Dynamic Non-IID Transfer Learning on Graphs
by: Wang, Haohui, et al.
Published: (2023)
by: Wang, Haohui, et al.
Published: (2023)
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
by: Moniri, Behrad, et al.
Published: (2023)
by: Moniri, Behrad, et al.
Published: (2023)
EVINET: Towards Open-World Graph Learning via Evidential Reasoning Network
by: Guan, Weijie, et al.
Published: (2025)
by: Guan, Weijie, et al.
Published: (2025)
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
by: Moniri, Behrad, et al.
Published: (2026)
by: Moniri, Behrad, et al.
Published: (2026)
Two-Layer Microwave Linear Analog Computer (MiLAC)-aided Multi-user MISO Networks
by: Zhou, Xiaohua, et al.
Published: (2026)
by: Zhou, Xiaohua, et al.
Published: (2026)
How Two-Layer Neural Networks Learn, One (Giant) Step at a Time
by: Dandi, Yatin, et al.
Published: (2023)
by: Dandi, Yatin, et al.
Published: (2023)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
by: Kothapalli, Vignesh, et al.
Published: (2025)
by: Kothapalli, Vignesh, et al.
Published: (2025)
Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias
by: Hu, Yuanzhe, et al.
Published: (2025)
by: Hu, Yuanzhe, et al.
Published: (2025)
A Model Zoo on Phase Transitions in Neural Networks
by: Schürholt, Konstantin, et al.
Published: (2025)
by: Schürholt, Konstantin, et al.
Published: (2025)
Towards Heterogeneous Long-tailed Learning: Benchmarking, Metrics, and Toolbox
by: Wang, Haohui, et al.
Published: (2023)
by: Wang, Haohui, et al.
Published: (2023)
New Evidence of the Two-Phase Learning Dynamics of Neural Networks
by: Zhou, Zhanpeng, et al.
Published: (2025)
by: Zhou, Zhanpeng, et al.
Published: (2025)
LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection
by: Zeng, Xinyue, et al.
Published: (2025)
by: Zeng, Xinyue, et al.
Published: (2025)
Divergence-free Linearized Neural Networks: Integral Representation and Optimal Approximation Rates
by: He, Juncai, et al.
Published: (2026)
by: He, Juncai, et al.
Published: (2026)
Structured and Balanced Multi-Component and Multi-Layer Neural Networks
by: Zhang, Shijun, et al.
Published: (2024)
by: Zhang, Shijun, et al.
Published: (2024)
Temporal Flexibility in Spiking Neural Networks: Towards Generalization Across Time Steps and Deployment Friendliness
by: Du, Kangrui, et al.
Published: (2025)
by: Du, Kangrui, et al.
Published: (2025)
Optimal Convergence Rates of Deep Neural Network Classifiers
by: Zhang, Zihan, et al.
Published: (2025)
by: Zhang, Zihan, et al.
Published: (2025)
WATS: Calibrating Graph Neural Networks with Wavelet-Aware Temperature Scaling
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
by: Dominé, Clémentine C. J., et al.
Published: (2024)
by: Dominé, Clémentine C. J., et al.
Published: (2024)
Global Universality of the Two-Layer Neural Network with thek-Rectified Linear Unit
by: Naoya Hatano, et al.
Published: (2024)
by: Naoya Hatano, et al.
Published: (2024)
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
by: Kothapalli, Vignesh, et al.
Published: (2026)
by: Kothapalli, Vignesh, et al.
Published: (2026)
Spectral Insights into Data-Oblivious Critical Layers in Large Language Models
by: Liu, Xuyuan, et al.
Published: (2025)
by: Liu, Xuyuan, et al.
Published: (2025)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
by: Kosson, Atli, et al.
Published: (2023)
by: Kosson, Atli, et al.
Published: (2023)
Optimal ISAC Beamforming Structure and Efficient Algorithms for Sum Rate and CRLB Balancing
by: Fang, Tianyu, et al.
Published: (2025)
by: Fang, Tianyu, et al.
Published: (2025)
GLinSAT: The General Linear Satisfiability Neural Network Layer By Accelerated Gradient Descent
by: Zeng, Hongtai, et al.
Published: (2024)
by: Zeng, Hongtai, et al.
Published: (2024)
Improving the Quality of Life for Older Adult Residents in Long‐Term Care Facilities: Experiences From COVID‐19 Pandemic and Future Preparedness
by: Haohui Deng, et al.
Published: (2025)
by: Haohui Deng, et al.
Published: (2025)
Some Improved Results on Fair and Balanced Graph Partitions
by: Viswanathan, Vignesh
Published: (2026)
by: Viswanathan, Vignesh
Published: (2026)
Optimal Learning Rate Schedule for Balancing Effort and Performance
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
A Three-regime Model of Network Pruning
by: Zhou, Yefan, et al.
Published: (2023)
by: Zhou, Yefan, et al.
Published: (2023)
Enhancing Size Generalization in Graph Neural Networks through Disentangled Representation Learning
by: Huang, Zheng, et al.
Published: (2024)
by: Huang, Zheng, et al.
Published: (2024)
Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval
by: Yang, Junkai, et al.
Published: (2026)
by: Yang, Junkai, et al.
Published: (2026)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
by: Zhang, Tongcheng, et al.
Published: (2026)
by: Zhang, Tongcheng, et al.
Published: (2026)
Similar Items
-
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
by: Kothapalli, Vignesh, et al.
Published: (2024) -
Suspicious Alignment of SGD: A Fine-Grained Step Size Condition Analysis
by: Deng, Shenyang, et al.
Published: (2026) -
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
by: Deng, Shenyang, et al.
Published: (2026) -
KCES: Training-Free Defense for Robust Graph Neural Networks via Kernel Complexity
by: Jia, Yaning, et al.
Published: (2025) -
Can Kernel Methods Explain How the Data Affects Neural Collapse?
by: Kothapalli, Vignesh, et al.
Published: (2024)