Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Jiarui, Huang, Wei, Zhang, Miao, Suzuki, Taiji, Nie, Liqiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
di: Jiang, Jiarui, et al.
Pubblicazione: (2024)
di: Jiang, Jiarui, et al.
Pubblicazione: (2024)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
di: Oh, Junsoo, et al.
Pubblicazione: (2025)
di: Oh, Junsoo, et al.
Pubblicazione: (2025)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
di: Li, Bingrui, et al.
Pubblicazione: (2024)
di: Li, Bingrui, et al.
Pubblicazione: (2024)
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
A Polynomial-time Algorithm for Online Sparse Linear Regression with Improved Regret Bound under Weaker Conditions
di: Li, Junfan, et al.
Pubblicazione: (2025)
di: Li, Junfan, et al.
Pubblicazione: (2025)
Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
di: Chen, Yilan, et al.
Pubblicazione: (2025)
di: Chen, Yilan, et al.
Pubblicazione: (2025)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
di: Watanabe, Chihiro, et al.
Pubblicazione: (2021)
di: Watanabe, Chihiro, et al.
Pubblicazione: (2021)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
di: Wakayama, Tomoya, et al.
Pubblicazione: (2025)
di: Wakayama, Tomoya, et al.
Pubblicazione: (2025)
Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression
di: Ding, Shihong, et al.
Pubblicazione: (2025)
di: Ding, Shihong, et al.
Pubblicazione: (2025)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
di: Kim, Juno, et al.
Pubblicazione: (2024)
di: Kim, Juno, et al.
Pubblicazione: (2024)
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
Transformers are Minimax Optimal Nonparametric In-Context Learners
di: Kim, Juno, et al.
Pubblicazione: (2024)
di: Kim, Juno, et al.
Pubblicazione: (2024)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
di: Nishikawa, Naoki, et al.
Pubblicazione: (2025)
di: Nishikawa, Naoki, et al.
Pubblicazione: (2025)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
di: Yamamoto, Naoya, et al.
Pubblicazione: (2025)
di: Yamamoto, Naoya, et al.
Pubblicazione: (2025)
Learning Curves of Stochastic Gradient Descent in Kernel Regression
di: Zhang, Haihan, et al.
Pubblicazione: (2025)
di: Zhang, Haihan, et al.
Pubblicazione: (2025)
Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
di: Cheng, Xiang, et al.
Pubblicazione: (2023)
di: Cheng, Xiang, et al.
Pubblicazione: (2023)
Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual Learning
di: Qiu, Haomiao, et al.
Pubblicazione: (2025)
di: Qiu, Haomiao, et al.
Pubblicazione: (2025)
Stochastic Gradient Descent for Nonparametric Additive Regression
di: Chen, Xin, et al.
Pubblicazione: (2024)
di: Chen, Xin, et al.
Pubblicazione: (2024)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
di: Chen, Zonghao, et al.
Pubblicazione: (2025)
di: Chen, Zonghao, et al.
Pubblicazione: (2025)
Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
di: Kim, Juno, et al.
Pubblicazione: (2025)
di: Kim, Juno, et al.
Pubblicazione: (2025)
Test time training enhances in-context learning of nonlinear functions
di: Kuwataka, Kento, et al.
Pubblicazione: (2025)
di: Kuwataka, Kento, et al.
Pubblicazione: (2025)
Deep Two-Way Matrix Reordering for Relational Data Analysis
di: Watanabe, Chihiro, et al.
Pubblicazione: (2021)
di: Watanabe, Chihiro, et al.
Pubblicazione: (2021)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
di: Awano, Ryoya, et al.
Pubblicazione: (2026)
di: Awano, Ryoya, et al.
Pubblicazione: (2026)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
di: Nishikawa, Naoki, et al.
Pubblicazione: (2024)
di: Nishikawa, Naoki, et al.
Pubblicazione: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
di: Kim, Juno, et al.
Pubblicazione: (2024)
di: Kim, Juno, et al.
Pubblicazione: (2024)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
di: Takakura, Shokichi, et al.
Pubblicazione: (2024)
di: Takakura, Shokichi, et al.
Pubblicazione: (2024)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
di: Kawata, Ryotaro, et al.
Pubblicazione: (2026)
di: Kawata, Ryotaro, et al.
Pubblicazione: (2026)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
di: Takakura, Shokichi, et al.
Pubblicazione: (2023)
di: Takakura, Shokichi, et al.
Pubblicazione: (2023)
Comparing the Moore-Penrose Pseudoinverse and Gradient Descent for Solving Linear Regression Problems: A Performance Analysis
di: Adams, Alex
Pubblicazione: (2025)
di: Adams, Alex
Pubblicazione: (2025)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
di: Bu, Dake, et al.
Pubblicazione: (2024)
di: Bu, Dake, et al.
Pubblicazione: (2024)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
di: Bu, Dake, et al.
Pubblicazione: (2025)
di: Bu, Dake, et al.
Pubblicazione: (2025)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
di: Zhang, Tongcheng, et al.
Pubblicazione: (2026)
Private Gradient Descent for Linear Regression: Tighter Error Bounds and Instance-Specific Uncertainty Estimation
di: Brown, Gavin, et al.
Pubblicazione: (2024)
di: Brown, Gavin, et al.
Pubblicazione: (2024)
Harmonized Gradient Descent for Class Imbalanced Data Stream Online Learning
di: Zhou, Han, et al.
Pubblicazione: (2025)
di: Zhou, Han, et al.
Pubblicazione: (2025)
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training
di: Zhao, Shen-Yi, et al.
Pubblicazione: (2020)
di: Zhao, Shen-Yi, et al.
Pubblicazione: (2020)
SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting
di: Qiu, Haomiao, et al.
Pubblicazione: (2025)
di: Qiu, Haomiao, et al.
Pubblicazione: (2025)
Understanding Gradient Descent through the Training Jacobian
di: Belrose, Nora, et al.
Pubblicazione: (2024)
di: Belrose, Nora, et al.
Pubblicazione: (2024)
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
di: Kawata, Ryotaro, et al.
Pubblicazione: (2025)
di: Kawata, Ryotaro, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
di: Jiang, Jiarui, et al.
Pubblicazione: (2024) -
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
di: Oh, Junsoo, et al.
Pubblicazione: (2025) -
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
di: Li, Bingrui, et al.
Pubblicazione: (2024) -
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
di: Huang, Wei, et al.
Pubblicazione: (2025) -
A Polynomial-time Algorithm for Online Sparse Linear Regression with Improved Regret Bound under Weaker Conditions
di: Li, Junfan, et al.
Pubblicazione: (2025)