Demystifying Lazy Training of Neural Networks from a Macroscopic Viewpoint
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuqing, Luo, Tao, Zhou, Qixuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying Training Difficulty and Accelerating Convergence in Neural Network-Based PDE Solvers
by: Chen, Chuqi, et al.
Published: (2024)
by: Chen, Chuqi, et al.
Published: (2024)
Demystifying Distributed Training of Graph Neural Networks for Link Prediction
by: Huang, Xin, et al.
Published: (2025)
by: Huang, Xin, et al.
Published: (2025)
A priori Estimates for Deep Residual Network in Continuous-time Reinforcement Learning
by: Yin, Shuyu, et al.
Published: (2024)
by: Yin, Shuyu, et al.
Published: (2024)
ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks
by: Wu, Haixu, et al.
Published: (2025)
by: Wu, Haixu, et al.
Published: (2025)
LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
by: Shen, Xuan, et al.
Published: (2024)
by: Shen, Xuan, et al.
Published: (2024)
Residual Attention Physics-Informed Neural Networks for Robust Multiphysics Simulation of Steady-State Electrothermal Energy Systems
by: Zhou, Yuqing, et al.
Published: (2026)
by: Zhou, Yuqing, et al.
Published: (2026)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
PRISM: Demystifying Retention and Interaction in Mid-Training
by: Runwal, Bharat, et al.
Published: (2026)
by: Runwal, Bharat, et al.
Published: (2026)
Demystifying Oversmoothing in Attention-Based Graph Neural Networks
by: Wu, Xinyi, et al.
Published: (2023)
by: Wu, Xinyi, et al.
Published: (2023)
Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime
by: Oberweis, Noah, et al.
Published: (2025)
by: Oberweis, Noah, et al.
Published: (2025)
Demystifying Higher-Order Graph Neural Networks
by: Besta, Maciej, et al.
Published: (2024)
by: Besta, Maciej, et al.
Published: (2024)
Has the Deep Neural Network learned the Stochastic Process? An Evaluation Viewpoint
by: Kumar, Harshit, et al.
Published: (2024)
by: Kumar, Harshit, et al.
Published: (2024)
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
by: Dominé, Clémentine C. J., et al.
Published: (2024)
by: Dominé, Clémentine C. J., et al.
Published: (2024)
AutoSGNN: Automatic Propagation Mechanism Discovery for Spectral Graph Neural Networks
by: Mo, Shibing, et al.
Published: (2024)
by: Mo, Shibing, et al.
Published: (2024)
Lazy FSCA for Unsupervised Variable Selection
by: Zocco, Federico, et al.
Published: (2021)
by: Zocco, Federico, et al.
Published: (2021)
Enhancing Trustworthiness of Graph Neural Networks with Rank-Based Conformal Training
by: Wang, Ting, et al.
Published: (2025)
by: Wang, Ting, et al.
Published: (2025)
Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes
by: Tu, Zhenfeng, et al.
Published: (2024)
by: Tu, Zhenfeng, et al.
Published: (2024)
Demystifying Network Foundation Models
by: Beltiukov, Sylee, et al.
Published: (2025)
by: Beltiukov, Sylee, et al.
Published: (2025)
DR-CircuitGNN: Training Acceleration of Heterogeneous Circuit Graph Neural Network on GPUs
by: Luo, Yuebo, et al.
Published: (2025)
by: Luo, Yuebo, et al.
Published: (2025)
APEX: Probing Neural Networks via Activation Perturbation
by: Ren, Tao, et al.
Published: (2026)
by: Ren, Tao, et al.
Published: (2026)
Partially Lazy Gradient Descent for Smoothed Online Learning
by: Mhaisen, Naram, et al.
Published: (2026)
by: Mhaisen, Naram, et al.
Published: (2026)
Phase Diagram of Initial Condensation for Two-layer Neural Networks
by: Chen, Zhengan, et al.
Published: (2023)
by: Chen, Zhengan, et al.
Published: (2023)
Loss Spike in Training Neural Networks
by: Li, Xiaolong, et al.
Published: (2023)
by: Li, Xiaolong, et al.
Published: (2023)
Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
by: Xu, Zhenghao, et al.
Published: (2024)
by: Xu, Zhenghao, et al.
Published: (2024)
Why Are DMD Students Lazy? Understanding the Copying Behavior in Few-Step Distillation
by: Li, Shucheng, et al.
Published: (2026)
by: Li, Shucheng, et al.
Published: (2026)
John Ellipsoids via Lazy Updates
by: Woodruff, David P., et al.
Published: (2025)
by: Woodruff, David P., et al.
Published: (2025)
Sharpened Lazy Incremental Quasi-Newton Method
by: Lahoti, Aakash, et al.
Published: (2023)
by: Lahoti, Aakash, et al.
Published: (2023)
BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
by: Zhou, Wenjie, et al.
Published: (2025)
by: Zhou, Wenjie, et al.
Published: (2025)
On the Dataless Training of Neural Networks
by: Velasquez, Alvaro, et al.
Published: (2025)
by: Velasquez, Alvaro, et al.
Published: (2025)
Demystify Protein Generation with Hierarchical Conditional Diffusion Models
by: Ling, Zinan, et al.
Published: (2025)
by: Ling, Zinan, et al.
Published: (2025)
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
by: He, Jianliang, et al.
Published: (2025)
by: He, Jianliang, et al.
Published: (2025)
Frequency-adaptive Multi-scale Deep Neural Networks
by: Huang, Jizu, et al.
Published: (2024)
by: Huang, Jizu, et al.
Published: (2024)
Flow Matching from Viewpoint of Proximal Operators
by: Fukumizu, Kenji, et al.
Published: (2026)
by: Fukumizu, Kenji, et al.
Published: (2026)
Demystifying Prediction Powered Inference
by: Song, Yilin, et al.
Published: (2026)
by: Song, Yilin, et al.
Published: (2026)
Uncovering Critical Sets of Deep Neural Networks via Sample-Independent Critical Lifting
by: Zhang, Leyang, et al.
Published: (2025)
by: Zhang, Leyang, et al.
Published: (2025)
Gradient Rewiring for Editable Graph Neural Network Training
by: Jiang, Zhimeng, et al.
Published: (2024)
by: Jiang, Zhimeng, et al.
Published: (2024)
Uncertainty in Graph Neural Networks: A Survey
by: Wang, Fangxin, et al.
Published: (2024)
by: Wang, Fangxin, et al.
Published: (2024)
LazyDP: Co-Designing Algorithm-Software for Scalable Training of Differentially Private Recommendation Models
by: Lim, Juntaek, et al.
Published: (2024)
by: Lim, Juntaek, et al.
Published: (2024)
Policy Compatible Skill Incremental Learning via Lazy Learning Interface
by: Lee, Daehee, et al.
Published: (2025)
by: Lee, Daehee, et al.
Published: (2025)
On Multi-Stage Loss Dynamics in Neural Networks: Mechanisms of Plateau and Descent Stages
by: Chen, Zheng-An, et al.
Published: (2024)
by: Chen, Zheng-An, et al.
Published: (2024)
Similar Items
-
Quantifying Training Difficulty and Accelerating Convergence in Neural Network-Based PDE Solvers
by: Chen, Chuqi, et al.
Published: (2024) -
Demystifying Distributed Training of Graph Neural Networks for Link Prediction
by: Huang, Xin, et al.
Published: (2025) -
A priori Estimates for Deep Residual Network in Continuous-time Reinforcement Learning
by: Yin, Shuyu, et al.
Published: (2024) -
ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks
by: Wu, Haixu, et al.
Published: (2025) -
LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
by: Shen, Xuan, et al.
Published: (2024)