Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
Fuente:
arXiv
Saved in:
| Main Authors: | Gatmiry, Khashayar, Saunshi, Nikunj, Reddi, Sashank J., Jegelka, Stefanie, Kumar, Sanjiv |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Reasoning with Latent Thoughts: On the Power of Looped Transformers
by: Saunshi, Nikunj, et al.
Published: (2025)
by: Saunshi, Nikunj, et al.
Published: (2025)
On the Inductive Bias of Stacking Towards Improving Reasoning
by: Saunshi, Nikunj, et al.
Published: (2024)
by: Saunshi, Nikunj, et al.
Published: (2024)
Simplicity Bias via Global Convergence of Sharpness Minimization
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Landscape-Aware Growing: The Power of a Little LAG
by: Karp, Stefani, et al.
Published: (2024)
by: Karp, Stefani, et al.
Published: (2024)
Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
by: Huang, Jianhao, et al.
Published: (2025)
by: Huang, Jianhao, et al.
Published: (2025)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
Computing Optimal Regularizers for Online Linear Optimization
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel
by: Cutler, Dylan, et al.
Published: (2025)
by: Cutler, Dylan, et al.
Published: (2025)
Near-Optimal Algorithms for Group Distributionally Robust Optimization and Beyond
by: Soma, Tasuku, et al.
Published: (2022)
by: Soma, Tasuku, et al.
Published: (2022)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
Adversarial Online Learning with Temporal Feedback Graphs
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Learning with Exact Invariances in Polynomial Time
by: Soleymani, Ashkan, et al.
Published: (2025)
by: Soleymani, Ashkan, et al.
Published: (2025)
Rethinking Invariance in In-context Learning
by: Fang, Lizhe, et al.
Published: (2025)
by: Fang, Lizhe, et al.
Published: (2025)
Elastic Multi-Gradient Descent for Parallel Continual Learning
by: Lyu, Fan, et al.
Published: (2024)
by: Lyu, Fan, et al.
Published: (2024)
Conflict-Averse Gradient Descent for Multi-task Learning
by: Liu, Bo, et al.
Published: (2021)
by: Liu, Bo, et al.
Published: (2021)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Efficient Document Ranking with Learnable Late Interactions
by: Ji, Ziwei, et al.
Published: (2024)
by: Ji, Ziwei, et al.
Published: (2024)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
by: Guo, Xiaojun, et al.
Published: (2025)
by: Guo, Xiaojun, et al.
Published: (2025)
Can LLMs predict the convergence of Stochastic Gradient Descent?
by: Zekri, Oussama, et al.
Published: (2024)
by: Zekri, Oussama, et al.
Published: (2024)
Geometrically Inspired Kernel Machines for Collaborative Learning Beyond Gradient Descent
by: Kumar, Mohit, et al.
Published: (2024)
by: Kumar, Mohit, et al.
Published: (2024)
A Unified Approach to Controlling Implicit Regularization via Mirror Descent
by: Sun, Haoyuan, et al.
Published: (2023)
by: Sun, Haoyuan, et al.
Published: (2023)
Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation
by: Lukasik, Michal, et al.
Published: (2025)
by: Lukasik, Michal, et al.
Published: (2025)
The Initialization Determines Whether In-Context Learning Is Gradient Descent
by: Xie, Shifeng, et al.
Published: (2025)
by: Xie, Shifeng, et al.
Published: (2025)
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
by: Garg, Ishir, et al.
Published: (2026)
by: Garg, Ishir, et al.
Published: (2026)
Descent-Guided Policy Gradient for Scalable Cooperative Multi-Agent Learning
by: Yang, Shan, et al.
Published: (2026)
by: Yang, Shan, et al.
Published: (2026)
Structured Preconditioners in Adaptive Optimization: A Unified Analysis
by: Xie, Shuo, et al.
Published: (2025)
by: Xie, Shuo, et al.
Published: (2025)
GradTree: Learning Axis-Aligned Decision Trees with Gradient Descent
by: Marton, Sascha, et al.
Published: (2023)
by: Marton, Sascha, et al.
Published: (2023)
Survey on Generalization Theory for Graph Neural Networks
by: Vasileiou, Antonis, et al.
Published: (2025)
by: Vasileiou, Antonis, et al.
Published: (2025)
Natural Mitigation of Catastrophic Interference: Continual Learning in Power-Law Learning Environments
by: Gandhi, Atith, et al.
Published: (2024)
by: Gandhi, Atith, et al.
Published: (2024)
LAuReL: Learned Augmented Residual Layer
by: Menghani, Gaurav, et al.
Published: (2024)
by: Menghani, Gaurav, et al.
Published: (2024)
Learning Mixtures of Gaussians Using Diffusion Models
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Gradient Descent Algorithm Survey
by: Fucheng, Deng, et al.
Published: (2025)
by: Fucheng, Deng, et al.
Published: (2025)
PSMGD: Periodic Stochastic Multi-Gradient Descent for Fast Multi-Objective Optimization
by: Xu, Mingjing, et al.
Published: (2024)
by: Xu, Mingjing, et al.
Published: (2024)
Can In-context Learning Really Generalize to Out-of-distribution Tasks?
by: Wang, Qixun, et al.
Published: (2024)
by: Wang, Qixun, et al.
Published: (2024)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
by: Arda, Enes, et al.
Published: (2026)
by: Arda, Enes, et al.
Published: (2026)
Similar Items
-
On the Role of Depth and Looping for In-Context Learning with Task Diversity
by: Gatmiry, Khashayar, et al.
Published: (2024) -
Reasoning with Latent Thoughts: On the Power of Looped Transformers
by: Saunshi, Nikunj, et al.
Published: (2025) -
On the Inductive Bias of Stacking Towards Improving Reasoning
by: Saunshi, Nikunj, et al.
Published: (2024) -
Simplicity Bias via Global Convergence of Sharpness Minimization
by: Gatmiry, Khashayar, et al.
Published: (2024) -
Landscape-Aware Growing: The Power of a Little LAG
by: Karp, Stefani, et al.
Published: (2024)