The Initialization Determines Whether In-Context Learning Is Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Shifeng, Yuan, Rui, Rossi, Simone, Hannagan, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024)
by: Abuduweili, Abulikemu, et al.
Published: (2024)
From predictions to confidence intervals: an empirical study of conformal prediction methods for in-context learning
by: Huang, Zhe, et al.
Published: (2025)
by: Huang, Zhe, et al.
Published: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Stochastic Gradient Descent with Momentum is Algorithmically Stable
by: Lei, Yunwen, et al.
Published: (2026)
by: Lei, Yunwen, et al.
Published: (2026)
Variational Graph Contrastive Learning
by: Xie, Shifeng, et al.
Published: (2024)
by: Xie, Shifeng, et al.
Published: (2024)
Elastic Multi-Gradient Descent for Parallel Continual Learning
by: Lyu, Fan, et al.
Published: (2024)
by: Lyu, Fan, et al.
Published: (2024)
Conflict-Averse Gradient Descent for Multi-task Learning
by: Liu, Bo, et al.
Published: (2021)
by: Liu, Bo, et al.
Published: (2021)
Gradient Descent Algorithm Survey
by: Fucheng, Deng, et al.
Published: (2025)
by: Fucheng, Deng, et al.
Published: (2025)
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
by: Garg, Ishir, et al.
Published: (2026)
by: Garg, Ishir, et al.
Published: (2026)
Randomness and Interpolation Improve Gradient Descent
by: Li, Jiawen, et al.
Published: (2025)
by: Li, Jiawen, et al.
Published: (2025)
ONG: Orthogonal Natural Gradient Descent
by: Yadav, Yajat, et al.
Published: (2025)
by: Yadav, Yajat, et al.
Published: (2025)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
by: Huang, Jianhao, et al.
Published: (2025)
by: Huang, Jianhao, et al.
Published: (2025)
Geometrically Inspired Kernel Machines for Collaborative Learning Beyond Gradient Descent
by: Kumar, Mohit, et al.
Published: (2024)
by: Kumar, Mohit, et al.
Published: (2024)
GradTree: Learning Axis-Aligned Decision Trees with Gradient Descent
by: Marton, Sascha, et al.
Published: (2023)
by: Marton, Sascha, et al.
Published: (2023)
Subgraph Gaussian Embedding Contrast for Self-Supervised Graph Representation Learning
by: Xie, Shifeng, et al.
Published: (2025)
by: Xie, Shifeng, et al.
Published: (2025)
Adaptive Heavy-Tailed Stochastic Gradient Descent
by: Gong, Bodu, et al.
Published: (2025)
by: Gong, Bodu, et al.
Published: (2025)
Vanilla Gradient Descent for Oblique Decision Trees
by: Panda, Subrat Prasad, et al.
Published: (2024)
by: Panda, Subrat Prasad, et al.
Published: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
by: Gatmiry, Khashayar, et al.
Published: (2024)
by: Gatmiry, Khashayar, et al.
Published: (2024)
Efficient Search for Customized Activation Functions with Gradient Descent
by: Strack, Lukas, et al.
Published: (2024)
by: Strack, Lukas, et al.
Published: (2024)
Noise Balance and Stationary Distribution of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2023)
by: Ziyin, Liu, et al.
Published: (2023)
Can LLMs predict the convergence of Stochastic Gradient Descent?
by: Zekri, Oussama, et al.
Published: (2024)
by: Zekri, Oussama, et al.
Published: (2024)
FedBCD:Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
by: Liu, Junkang, et al.
Published: (2026)
by: Liu, Junkang, et al.
Published: (2026)
Gradient Descent Efficiency Index
by: Dhingra, Aviral
Published: (2024)
by: Dhingra, Aviral
Published: (2024)
Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training
by: Ma, Yuhan, et al.
Published: (2024)
by: Ma, Yuhan, et al.
Published: (2024)
Geodesic Gradient Descent: A Generic and Learning-rate-free Optimizer on Objective Function-induced Manifolds
by: Hu, Liwei, et al.
Published: (2026)
by: Hu, Liwei, et al.
Published: (2026)
Stochastic Re-weighted Gradient Descent via Distributionally Robust Optimization
by: Kumar, Ramnath, et al.
Published: (2023)
by: Kumar, Ramnath, et al.
Published: (2023)
Natural Gradient Descent for Online Continual Learning
by: Khawand, Joe, et al.
Published: (2026)
by: Khawand, Joe, et al.
Published: (2026)
Trustworthiness of Stochastic Gradient Descent in Distributed Learning
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
Descent-Guided Policy Gradient for Scalable Cooperative Multi-Agent Learning
by: Yang, Shan, et al.
Published: (2026)
by: Yang, Shan, et al.
Published: (2026)
Adaptive Online Mirror Descent for Tchebycheff Scalarization in Multi-Objective Learning
by: Liu, Meitong, et al.
Published: (2024)
by: Liu, Meitong, et al.
Published: (2024)
Compact Rule-Based Classifier Learning via Gradient Descent
by: Fumanal-Idocin, Javier, et al.
Published: (2025)
by: Fumanal-Idocin, Javier, et al.
Published: (2025)
PSMGD: Periodic Stochastic Multi-Gradient Descent for Fast Multi-Objective Optimization
by: Xu, Mingjing, et al.
Published: (2024)
by: Xu, Mingjing, et al.
Published: (2024)
Optimization, Generalization and Differential Privacy Bounds for Gradient Descent on Kolmogorov-Arnold Networks
by: Wang, Puyu, et al.
Published: (2026)
by: Wang, Puyu, et al.
Published: (2026)
Reconstructing Deep Neural Networks: Unleashing the Optimization Potential of Natural Gradient Descent
by: Liu, Weihua, et al.
Published: (2024)
by: Liu, Weihua, et al.
Published: (2024)
Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order Optimization
by: Liang, Chen, et al.
Published: (2026)
by: Liang, Chen, et al.
Published: (2026)
Generalized Euler Logarithm and its Applications in Machine Learning: Natural Gradient, Backpropagation, Generalized EG, Mirror Descent and OLPS
by: Cichocki, Andrzej
Published: (2025)
by: Cichocki, Andrzej
Published: (2025)
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
by: Lu, Yishun, et al.
Published: (2025)
by: Lu, Yishun, et al.
Published: (2025)
Auto-Unrolled Proximal Gradient Descent: An AutoML Approach to Interpretable Waveform Optimization
by: Kaplan, Ahmet
Published: (2026)
by: Kaplan, Ahmet
Published: (2026)
Similar Items
-
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024) -
From predictions to confidence intervals: an empirical study of conformal prediction methods for in-context learning
by: Huang, Zhe, et al.
Published: (2025) -
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023) -
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024) -
Stochastic Gradient Descent with Momentum is Algorithmically Stable
by: Lei, Yunwen, et al.
Published: (2026)