On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Tongcheng, Zhou, Zhanpeng, Wang, Mingze, Han, Andi, Huang, Wei, Suzuki, Taiji, Yan, Junchi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Role of Label Noise in the Feature Learning Process
por: Han, Andi, et al.
Publicado: (2025)
por: Han, Andi, et al.
Publicado: (2025)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
por: Huang, Wei, et al.
Publicado: (2026)
por: Huang, Wei, et al.
Publicado: (2026)
On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm
por: Zhou, Zhanpeng, et al.
Publicado: (2024)
por: Zhou, Zhanpeng, et al.
Publicado: (2024)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
por: Li, Bingrui, et al.
Publicado: (2024)
por: Li, Bingrui, et al.
Publicado: (2024)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
por: Wang, Jinbo, et al.
Publicado: (2025)
por: Wang, Jinbo, et al.
Publicado: (2025)
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
por: Bu, Dake, et al.
Publicado: (2026)
por: Bu, Dake, et al.
Publicado: (2026)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
por: Bu, Dake, et al.
Publicado: (2025)
por: Bu, Dake, et al.
Publicado: (2025)
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
por: Bu, Dake, et al.
Publicado: (2025)
por: Bu, Dake, et al.
Publicado: (2025)
New Evidence of the Two-Phase Learning Dynamics of Neural Networks
por: Zhou, Zhanpeng, et al.
Publicado: (2025)
por: Zhou, Zhanpeng, et al.
Publicado: (2025)
Learning to Solve Combinatorial Optimization under Positive Linear Constraints via Non-Autoregressive Neural Networks
por: Wang, Runzhong, et al.
Publicado: (2024)
por: Wang, Runzhong, et al.
Publicado: (2024)
Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise
por: Bae, Youngkyoung, et al.
Publicado: (2024)
por: Bae, Youngkyoung, et al.
Publicado: (2024)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
por: Kim, Juno, et al.
Publicado: (2025)
por: Kim, Juno, et al.
Publicado: (2025)
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
por: Huang, Wei, et al.
Publicado: (2025)
por: Huang, Wei, et al.
Publicado: (2025)
Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization
por: Zhang, Kuan, et al.
Publicado: (2025)
por: Zhang, Kuan, et al.
Publicado: (2025)
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
por: Chen, Feng, et al.
Publicado: (2023)
por: Chen, Feng, et al.
Publicado: (2023)
Hide and Seek in Noise Labels: Noise-Robust Collaborative Active Learning with LLM-Powered Assistance
por: Yuan, Bo, et al.
Publicado: (2025)
por: Yuan, Bo, et al.
Publicado: (2025)
How Transformers Learn to Plan via Multi-Token Prediction
por: Huang, Jianhao, et al.
Publicado: (2026)
por: Huang, Jianhao, et al.
Publicado: (2026)
SE-Merging: A Self-Enhanced Approach for Dynamic Model Merging
por: Chen, Zijun, et al.
Publicado: (2025)
por: Chen, Zijun, et al.
Publicado: (2025)
SGFormer: Single-Layer Graph Transformers with Approximation-Free Linear Complexity
por: Wu, Qitian, et al.
Publicado: (2024)
por: Wu, Qitian, et al.
Publicado: (2024)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
por: Kim, Jihwan, et al.
Publicado: (2026)
por: Kim, Jihwan, et al.
Publicado: (2026)
Learning under Temporal Label Noise
por: Nagaraj, Sujay, et al.
Publicado: (2024)
por: Nagaraj, Sujay, et al.
Publicado: (2024)
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
por: Wei, Chengkun, et al.
Publicado: (2025)
por: Wei, Chengkun, et al.
Publicado: (2025)
FedFixer: Mitigating Heterogeneous Label Noise in Federated Learning
por: Ji, Xinyuan, et al.
Publicado: (2024)
por: Ji, Xinyuan, et al.
Publicado: (2024)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
por: Takakura, Shokichi, et al.
Publicado: (2024)
por: Takakura, Shokichi, et al.
Publicado: (2024)
Minibatch and Local SGD: Algorithmic Stability and Linear Speedup in Generalization
por: Lei, Yunwen, et al.
Publicado: (2023)
por: Lei, Yunwen, et al.
Publicado: (2023)
Impact of Label Noise on Learning Complex Features
por: Vashisht, Rahul, et al.
Publicado: (2024)
por: Vashisht, Rahul, et al.
Publicado: (2024)
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
por: Zhang, Hanling, et al.
Publicado: (2025)
por: Zhang, Hanling, et al.
Publicado: (2025)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
por: Wachi, Akifumi, et al.
Publicado: (2026)
por: Wachi, Akifumi, et al.
Publicado: (2026)
Robust Deep Hawkes Process under Label Noise of Both Event and Occurrence
por: Tan, Xiaoyu, et al.
Publicado: (2024)
por: Tan, Xiaoyu, et al.
Publicado: (2024)
Resurrecting Label Propagation for Graphs with Heterophily and Label Noise
por: Cheng, Yao, et al.
Publicado: (2023)
por: Cheng, Yao, et al.
Publicado: (2023)
The Malignant Tail: Spectral Segregation of Label Noise in Over-Parameterized Networks
por: Wang, Zice
Publicado: (2026)
por: Wang, Zice
Publicado: (2026)
PTCL: Pseudo-Label Temporal Curriculum Learning for Label-Limited Dynamic Graph
por: Zhang, Shengtao, et al.
Publicado: (2025)
por: Zhang, Shengtao, et al.
Publicado: (2025)
Learning Divergence Fields for Shift-Robust Graph Representations
por: Wu, Qitian, et al.
Publicado: (2024)
por: Wu, Qitian, et al.
Publicado: (2024)
Addressing Long-Tail Noisy Label Learning Problems: a Two-Stage Solution with Label Refurbishment Considering Label Rarity
por: Wu, Ying-Hsuan, et al.
Publicado: (2024)
por: Wu, Ying-Hsuan, et al.
Publicado: (2024)
Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late in Training
por: Zhou, Zhanpeng, et al.
Publicado: (2024)
por: Zhou, Zhanpeng, et al.
Publicado: (2024)
NTKMTL: Mitigating Task Imbalance in Multi-Task Learning from Neural Tangent Kernel Perspective
por: Qin, Xiaohan, et al.
Publicado: (2025)
por: Qin, Xiaohan, et al.
Publicado: (2025)
RQP-SGD: Differential Private Machine Learning through Noisy SGD and Randomized Quantization
por: Feng, Ce, et al.
Publicado: (2024)
por: Feng, Ce, et al.
Publicado: (2024)
GSINA: Improving Subgraph Extraction for Graph Invariant Learning via Graph Sinkhorn Attention
por: Yan, Junchi, et al.
Publicado: (2024)
por: Yan, Junchi, et al.
Publicado: (2024)
MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training
por: Chen, Zhanpeng, et al.
Publicado: (2024)
por: Chen, Zhanpeng, et al.
Publicado: (2024)
Transformers from Diffusion: A Unified Framework for Neural Message Passing
por: Wu, Qitian, et al.
Publicado: (2024)
por: Wu, Qitian, et al.
Publicado: (2024)
Ejemplares similares
-
On the Role of Label Noise in the Feature Learning Process
por: Han, Andi, et al.
Publicado: (2025) -
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
por: Huang, Wei, et al.
Publicado: (2026) -
On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm
por: Zhou, Zhanpeng, et al.
Publicado: (2024) -
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
por: Li, Bingrui, et al.
Publicado: (2024) -
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
por: Wang, Jinbo, et al.
Publicado: (2025)