Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
Fuente:
arXiv
Salvato in:
| Autori principali: | Cheng, Xiang, Chen, Yuxin, Sra, Suvrit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Graph Transformers Dream of Electric Flow
di: Cheng, Xiang, et al.
Pubblicazione: (2024)
di: Cheng, Xiang, et al.
Pubblicazione: (2024)
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
Efficient Sampling on Riemannian Manifolds via Langevin MCMC
di: Cheng, Xiang, et al.
Pubblicazione: (2024)
di: Cheng, Xiang, et al.
Pubblicazione: (2024)
Riemannian Bilevel Optimization
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part I
di: Tian, Yi, et al.
Pubblicazione: (2022)
di: Tian, Yi, et al.
Pubblicazione: (2022)
Cost-Driven Representation Learning for Linear Quadratic Gaussian Control: Part II
di: Tian, Yi, et al.
Pubblicazione: (2026)
di: Tian, Yi, et al.
Pubblicazione: (2026)
Linear attention is (maybe) all you need (to understand transformer optimization)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
Trees to Flows and Back: Unifying Decision Trees and Diffusion Models
di: Ramachandran, Sai Niranjan, et al.
Pubblicazione: (2026)
di: Ramachandran, Sai Niranjan, et al.
Pubblicazione: (2026)
Implicit Bias in Matrix Factorization and its Explicit Realization in a New Architecture
di: Hou, Yikun, et al.
Pubblicazione: (2025)
di: Hou, Yikun, et al.
Pubblicazione: (2025)
First-Order Methods for Linearly Constrained Bilevel Optimization
di: Kornowski, Guy, et al.
Pubblicazione: (2024)
di: Kornowski, Guy, et al.
Pubblicazione: (2024)
How to escape sharp minima with random perturbations
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
Distributed Gradient Descent for Functional Learning
di: Yu, Zhan, et al.
Pubblicazione: (2023)
di: Yu, Zhan, et al.
Pubblicazione: (2023)
A projection-based framework for gradient-free and parallel learning
di: Bergmeister, Andreas, et al.
Pubblicazione: (2025)
di: Bergmeister, Andreas, et al.
Pubblicazione: (2025)
Continuum Transformers Perform In-Context Learning by Operator Gradient Descent
di: Mishra, Abhiti, et al.
Pubblicazione: (2025)
di: Mishra, Abhiti, et al.
Pubblicazione: (2025)
Cross-fluctuation phase transitions reveal sampling dynamics in diffusion models
di: Ramachandran, Sai Niranjan, et al.
Pubblicazione: (2025)
di: Ramachandran, Sai Niranjan, et al.
Pubblicazione: (2025)
Tight Generalization Bounds for Noiseless Inverse Optimization
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
di: Fatemi, Pouria, et al.
Pubblicazione: (2026)
Revisiting Frank-Wolfe for Structured Nonconvex Optimization
di: Maskan, Hoomaan, et al.
Pubblicazione: (2025)
di: Maskan, Hoomaan, et al.
Pubblicazione: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
di: Shen, Lingfeng, et al.
Pubblicazione: (2023)
di: Shen, Lingfeng, et al.
Pubblicazione: (2023)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
di: Huang, Jianhao, et al.
Pubblicazione: (2025)
di: Huang, Jianhao, et al.
Pubblicazione: (2025)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
di: Jiang, Jiarui, et al.
Pubblicazione: (2025)
di: Jiang, Jiarui, et al.
Pubblicazione: (2025)
Gradient Descent Fails to Learn High-frequency Functions and Modular Arithmetic
di: Takhanov, Rustem, et al.
Pubblicazione: (2023)
di: Takhanov, Rustem, et al.
Pubblicazione: (2023)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
di: Gatmiry, Khashayar, et al.
Pubblicazione: (2024)
di: Gatmiry, Khashayar, et al.
Pubblicazione: (2024)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
Learning High-Dimensional Parity Functions with Product Networks using Gradient Descent
di: Larue, Guillaume, et al.
Pubblicazione: (2026)
di: Larue, Guillaume, et al.
Pubblicazione: (2026)
The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
di: Dandi, Yatin, et al.
Pubblicazione: (2025)
di: Dandi, Yatin, et al.
Pubblicazione: (2025)
Geodesic Gradient Descent: A Generic and Learning-rate-free Optimizer on Objective Function-induced Manifolds
di: Hu, Liwei, et al.
Pubblicazione: (2026)
di: Hu, Liwei, et al.
Pubblicazione: (2026)
One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning
di: He, Bowen, et al.
Pubblicazione: (2026)
di: He, Bowen, et al.
Pubblicazione: (2026)
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
Anytime Acceleration of Gradient Descent
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
Efficient Search for Customized Activation Functions with Gradient Descent
di: Strack, Lukas, et al.
Pubblicazione: (2024)
di: Strack, Lukas, et al.
Pubblicazione: (2024)
Functional Central Limit Theorem for Stochastic Gradient Descent
di: Flamand, Kessang, et al.
Pubblicazione: (2026)
di: Flamand, Kessang, et al.
Pubblicazione: (2026)
The Initialization Determines Whether In-Context Learning Is Gradient Descent
di: Xie, Shifeng, et al.
Pubblicazione: (2025)
di: Xie, Shifeng, et al.
Pubblicazione: (2025)
Understanding Task Vectors in In-Context Learning: Emergence, Functionality, and Limitations
di: Dong, Yuxin, et al.
Pubblicazione: (2025)
di: Dong, Yuxin, et al.
Pubblicazione: (2025)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
di: Aksenov, Yaroslav, et al.
Pubblicazione: (2024)
di: Aksenov, Yaroslav, et al.
Pubblicazione: (2024)
How Transformers Learn Causal Structure with Gradient Descent
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
di: Qin, Zhen, et al.
Pubblicazione: (2025)
di: Qin, Zhen, et al.
Pubblicazione: (2025)
Curl Descent: Non-Gradient Learning Dynamics with Sign-Diverse Plasticity
di: Ninou, Hugo, et al.
Pubblicazione: (2025)
di: Ninou, Hugo, et al.
Pubblicazione: (2025)
Unraveling the Gradient Descent Dynamics of Transformers
di: Song, Bingqing, et al.
Pubblicazione: (2024)
di: Song, Bingqing, et al.
Pubblicazione: (2024)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
di: Li, Bingrui, et al.
Pubblicazione: (2024)
di: Li, Bingrui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Graph Transformers Dream of Electric Flow
di: Cheng, Xiang, et al.
Pubblicazione: (2024) -
Linearly Convergent Algorithms for Nonsmooth Problems with Unknown Smooth Pieces
di: Zhang, Zhe, et al.
Pubblicazione: (2025) -
Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024) -
Efficient Sampling on Riemannian Manifolds via Langevin MCMC
di: Cheng, Xiang, et al.
Pubblicazione: (2024) -
Riemannian Bilevel Optimization
di: Dutta, Sanchayan, et al.
Pubblicazione: (2024)