On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Bingrui, Huang, Wei, Han, Andi, Zhou, Zhanpeng, Suzuki, Taiji, Zhu, Jun, Chen, Jianfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
von: Zhang, Tongcheng, et al.
Veröffentlicht: (2026)
von: Zhang, Tongcheng, et al.
Veröffentlicht: (2026)
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Efficient Hyperparameter Tuning via Trajectory Invariance Principle
von: Li, Bingrui, et al.
Veröffentlicht: (2025)
von: Li, Bingrui, et al.
Veröffentlicht: (2025)
Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
von: Chen, Yilan, et al.
Veröffentlicht: (2025)
von: Chen, Yilan, et al.
Veröffentlicht: (2025)
On the Role of Label Noise in the Feature Learning Process
von: Han, Andi, et al.
Veröffentlicht: (2025)
von: Han, Andi, et al.
Veröffentlicht: (2025)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
von: Jiang, Jiarui, et al.
Veröffentlicht: (2025)
von: Jiang, Jiarui, et al.
Veröffentlicht: (2025)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
von: Takakura, Shokichi, et al.
Veröffentlicht: (2024)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2024)
On the Comparison between Multi-modal and Single-modal Contrastive Learning
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
von: Bu, Dake, et al.
Veröffentlicht: (2024)
von: Bu, Dake, et al.
Veröffentlicht: (2024)
Stochastic Gradient Descent for Two-layer Neural Networks
von: Cao, Dinghao, et al.
Veröffentlicht: (2024)
von: Cao, Dinghao, et al.
Veröffentlicht: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
von: Huang, Wei, et al.
Veröffentlicht: (2026)
von: Huang, Wei, et al.
Veröffentlicht: (2026)
Transformers Provably Solve Parity Efficiently with Chain of Thought
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Quantifying the Optimization and Generalization Advantages of Graph Neural Networks Over Multilayer Perceptrons
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late in Training
von: Zhou, Zhanpeng, et al.
Veröffentlicht: (2024)
von: Zhou, Zhanpeng, et al.
Veröffentlicht: (2024)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
von: Awano, Ryoya, et al.
Veröffentlicht: (2026)
von: Awano, Ryoya, et al.
Veröffentlicht: (2026)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
von: Jiang, Jiarui, et al.
Veröffentlicht: (2024)
von: Jiang, Jiarui, et al.
Veröffentlicht: (2024)
Transformers are Minimax Optimal Nonparametric In-Context Learners
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
von: Bu, Dake, et al.
Veröffentlicht: (2026)
von: Bu, Dake, et al.
Veröffentlicht: (2026)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
von: Yamamoto, Naoya, et al.
Veröffentlicht: (2025)
von: Yamamoto, Naoya, et al.
Veröffentlicht: (2025)
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
von: Qin, Zhen, et al.
Veröffentlicht: (2025)
von: Qin, Zhen, et al.
Veröffentlicht: (2025)
Convergence Error Analysis of Reflected Gradient Langevin Dynamics for Globally Optimizing Non-Convex Constrained Problems
von: Sato, Kanji, et al.
Veröffentlicht: (2022)
von: Sato, Kanji, et al.
Veröffentlicht: (2022)
Unraveling the Gradient Descent Dynamics of Transformers
von: Song, Bingqing, et al.
Veröffentlicht: (2024)
von: Song, Bingqing, et al.
Veröffentlicht: (2024)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
von: Cheng, Wenhua, et al.
Veröffentlicht: (2023)
von: Cheng, Wenhua, et al.
Veröffentlicht: (2023)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
Test time training enhances in-context learning of nonlinear functions
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
Distributed Gradient Descent for Functional Learning
von: Yu, Zhan, et al.
Veröffentlicht: (2023)
von: Yu, Zhan, et al.
Veröffentlicht: (2023)
Curl Descent: Non-Gradient Learning Dynamics with Sign-Diverse Plasticity
von: Ninou, Hugo, et al.
Veröffentlicht: (2025)
von: Ninou, Hugo, et al.
Veröffentlicht: (2025)
Gradient Descent, Stochastic Optimization, and Other Tales
von: Lu, Jun
Veröffentlicht: (2022)
von: Lu, Jun
Veröffentlicht: (2022)
Ähnliche Einträge
-
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
von: Zhang, Tongcheng, et al.
Veröffentlicht: (2026) -
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
von: Huang, Wei, et al.
Veröffentlicht: (2025) -
Efficient Hyperparameter Tuning via Trajectory Invariance Principle
von: Li, Bingrui, et al.
Veröffentlicht: (2025) -
Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
von: Chen, Yilan, et al.
Veröffentlicht: (2025) -
On the Role of Label Noise in the Feature Learning Process
von: Han, Andi, et al.
Veröffentlicht: (2025)