New Evidence of the Two-Phase Learning Dynamics of Neural Networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Zhanpeng, Yang, Yongyi, Sugiyama, Mahito, Yan, Junchi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the Cone Effect in the Learning Dynamics
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2025)
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2025)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
Linear Mode Connectivity in Differentiable Tree Ensembles
di: Kanoh, Ryuichi, et al.
Pubblicazione: (2024)
di: Kanoh, Ryuichi, et al.
Pubblicazione: (2024)
A Complete Decomposition of KL Error using Refined Information and Mode Interaction Selection
di: Enouen, James, et al.
Pubblicazione: (2024)
di: Enouen, James, et al.
Pubblicazione: (2024)
When Graph Language Models Go Beyond Memorization
di: Yamada, Masatsugu, et al.
Pubblicazione: (2026)
di: Yamada, Masatsugu, et al.
Pubblicazione: (2026)
Pseudo-Nonlinear Data Augmentation: A Constrained Energy Minimization Viewpoint
di: Hu, Pingbang, et al.
Pubblicazione: (2024)
di: Hu, Pingbang, et al.
Pubblicazione: (2024)
StiefelGen: A Simple, Model Agnostic Approach for Time Series Data Augmentation over Riemannian Manifolds
di: Cheema, Prasad, et al.
Pubblicazione: (2024)
di: Cheema, Prasad, et al.
Pubblicazione: (2024)
Dual Riemannian Newton Method on Statistical Manifolds
di: Zhou, Derun, et al.
Pubblicazione: (2025)
di: Zhou, Derun, et al.
Pubblicazione: (2025)
How Graph Neural Networks Learn: Lessons from Training Dynamics
di: Yang, Chenxiao, et al.
Pubblicazione: (2023)
di: Yang, Chenxiao, et al.
Pubblicazione: (2023)
Bringing Structure to Naturalness: On the Naturalness of ASTs
di: Pârţachi, Profir-Petru, et al.
Pubblicazione: (2025)
di: Pârţachi, Profir-Petru, et al.
Pubblicazione: (2025)
Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late in Training
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2024)
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2024)
An Equivariance Toolbox for Learning Dynamics
di: Yang, Yongyi, et al.
Pubblicazione: (2025)
di: Yang, Yongyi, et al.
Pubblicazione: (2025)
On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2024)
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2024)
Learning to Solve Combinatorial Optimization under Positive Linear Constraints via Non-Autoregressive Neural Networks
di: Wang, Runzhong, et al.
Pubblicazione: (2024)
di: Wang, Runzhong, et al.
Pubblicazione: (2024)
Quadratic polarity and polar Fenchel-Young divergences from the canonical Legendre polarity
di: Nielsen, Frank, et al.
Pubblicazione: (2026)
di: Nielsen, Frank, et al.
Pubblicazione: (2026)
On the Role of Label Noise in the Feature Learning Process
di: Han, Andi, et al.
Pubblicazione: (2025)
di: Han, Andi, et al.
Pubblicazione: (2025)
HERTA: A High-Efficiency and Rigorous Training Algorithm for Unfolded Graph Neural Networks
di: Yang, Yongyi, et al.
Pubblicazione: (2024)
di: Yang, Yongyi, et al.
Pubblicazione: (2024)
Towards Quantum Graph Neural Networks: An Ego-Graph Learning Approach
di: Ai, Xing, et al.
Pubblicazione: (2022)
di: Ai, Xing, et al.
Pubblicazione: (2022)
EasyDGL: Encode, Train and Interpret for Continuous-time Dynamic Graph Learning
di: Chen, Chao, et al.
Pubblicazione: (2023)
di: Chen, Chao, et al.
Pubblicazione: (2023)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
di: Wang, Jinbo, et al.
Pubblicazione: (2025)
di: Wang, Jinbo, et al.
Pubblicazione: (2025)
Implicit vs Unfolded Graph Neural Networks
di: Yang, Yongyi, et al.
Pubblicazione: (2021)
di: Yang, Yongyi, et al.
Pubblicazione: (2021)
Same Graph, Different Likelihoods: Calibration of Autoregressive Graph Generators via Permutation-Equivalent Encodings
di: Fredsgaard, Laurits, et al.
Pubblicazione: (2026)
di: Fredsgaard, Laurits, et al.
Pubblicazione: (2026)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
di: Pang, Tianyu, et al.
Pubblicazione: (2026)
di: Pang, Tianyu, et al.
Pubblicazione: (2026)
Learning with Incomplete Context: Linear Contextual Bandits with Pretrained Imputation
di: Yan, Hao, et al.
Pubblicazione: (2025)
di: Yan, Hao, et al.
Pubblicazione: (2025)
NTKMTL: Mitigating Task Imbalance in Multi-Task Learning from Neural Tangent Kernel Perspective
di: Qin, Xiaohan, et al.
Pubblicazione: (2025)
di: Qin, Xiaohan, et al.
Pubblicazione: (2025)
Transformers from Diffusion: A Unified Framework for Neural Message Passing
di: Wu, Qitian, et al.
Pubblicazione: (2024)
di: Wu, Qitian, et al.
Pubblicazione: (2024)
Provable Low-Frequency Bias of In-Context Learning of Representations
di: Yang, Yongyi, et al.
Pubblicazione: (2025)
di: Yang, Yongyi, et al.
Pubblicazione: (2025)
A Model Zoo on Phase Transitions in Neural Networks
di: Schürholt, Konstantin, et al.
Pubblicazione: (2025)
di: Schürholt, Konstantin, et al.
Pubblicazione: (2025)
Swing-by Dynamics in Concept Learning and Compositional Generalization
di: Yang, Yongyi, et al.
Pubblicazione: (2024)
di: Yang, Yongyi, et al.
Pubblicazione: (2024)
Two Facets of SDE Under an Information-Theoretic Lens: Generalization of SGD via Training Trajectories and via Terminal States
di: Wang, Ziqiao, et al.
Pubblicazione: (2022)
di: Wang, Ziqiao, et al.
Pubblicazione: (2022)
Enhancing Size Generalization in Graph Neural Networks through Disentangled Representation Learning
di: Huang, Zheng, et al.
Pubblicazione: (2024)
di: Huang, Zheng, et al.
Pubblicazione: (2024)
Learning Divergence Fields for Shift-Robust Graph Representations
di: Wu, Qitian, et al.
Pubblicazione: (2024)
di: Wu, Qitian, et al.
Pubblicazione: (2024)
Topological Invariance and Breakdown in Learning
di: Yang, Yongyi, et al.
Pubblicazione: (2025)
di: Yang, Yongyi, et al.
Pubblicazione: (2025)
Euphonium: Steering Video Flow Matching via Process Reward Gradient Guided Stochastic Dynamics
di: Zhong, Ruizhe, et al.
Pubblicazione: (2026)
di: Zhong, Ruizhe, et al.
Pubblicazione: (2026)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
di: Li, Bingrui, et al.
Pubblicazione: (2024)
di: Li, Bingrui, et al.
Pubblicazione: (2024)
Non-stationary Online Learning for Curved Losses: Improved Dynamic Regret via Mixability
di: Zhang, Yu-Jie, et al.
Pubblicazione: (2025)
di: Zhang, Yu-Jie, et al.
Pubblicazione: (2025)
Molecule Generation for Drug Design: a Graph Learning Perspective
di: Yang, Nianzu, et al.
Pubblicazione: (2022)
di: Yang, Nianzu, et al.
Pubblicazione: (2022)
mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
di: Yang, Yongyi, et al.
Pubblicazione: (2026)
di: Yang, Yongyi, et al.
Pubblicazione: (2026)
SmartMixed: A Two-Phase Training Strategy for Adaptive Activation Function Learning in Neural Networks
di: Omidvar, Amin
Pubblicazione: (2025)
di: Omidvar, Amin
Pubblicazione: (2025)
How Does Gradient Descent Learn Features -- A Local Analysis for Regularized Two-Layer Neural Networks
di: Zhou, Mo, et al.
Pubblicazione: (2024)
di: Zhou, Mo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
On the Cone Effect in the Learning Dynamics
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2025) -
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026) -
Linear Mode Connectivity in Differentiable Tree Ensembles
di: Kanoh, Ryuichi, et al.
Pubblicazione: (2024) -
A Complete Decomposition of KL Error using Refined Information and Mode Interaction Selection
di: Enouen, James, et al.
Pubblicazione: (2024) -
When Graph Language Models Go Beyond Memorization
di: Yamada, Masatsugu, et al.
Pubblicazione: (2026)