Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Tong, Huang, Yu, Liang, Yingbin, Chi, Yuejie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Accelerating Convergence of Score-Based Diffusion Models, Provably
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
On the optimization dynamics of RLVR: Gradient gap and step size thresholds
by: Suk, Joe, et al.
Published: (2025)
by: Suk, Joe, et al.
Published: (2025)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Multi-beam Beamforming in RIS-aided MIMO Subject to Reradiation Mask Constraints -- Optimization and Machine Learning Design
by: Wang, Shumin, et al.
Published: (2025)
by: Wang, Shumin, et al.
Published: (2025)
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
by: Huang, Yu, et al.
Published: (2024)
by: Huang, Yu, et al.
Published: (2024)
Statistical and Algorithmic Foundations of Reinforcement Learning
by: Chi, Yuejie, et al.
Published: (2025)
by: Chi, Yuejie, et al.
Published: (2025)
Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
by: Li, Gen, et al.
Published: (2020)
by: Li, Gen, et al.
Published: (2020)
Provable Phase Retrieval with Mirror Descent
by: Godeme, Jean-Jacques, et al.
Published: (2022)
by: Godeme, Jean-Jacques, et al.
Published: (2022)
Structured Gradient Descent for Fast Robust Low-Rank Hankel Matrix Completion
by: Cai, HanQin, et al.
Published: (2022)
by: Cai, HanQin, et al.
Published: (2022)
Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization
by: Shen, Wei, et al.
Published: (2023)
by: Shen, Wei, et al.
Published: (2023)
Preconditioning Benefits of Spectral Orthogonalization in Muon
by: Ma, Jianhao, et al.
Published: (2026)
by: Ma, Jianhao, et al.
Published: (2026)
Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis
by: Li, Gen, et al.
Published: (2021)
by: Li, Gen, et al.
Published: (2021)
A Fast and Provable Algorithm for Sparse Phase Retrieval
by: Cai, Jian-Feng, et al.
Published: (2023)
by: Cai, Jian-Feng, et al.
Published: (2023)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
by: Arda, Enes, et al.
Published: (2026)
by: Arda, Enes, et al.
Published: (2026)
A Retention-Centric Framework for Continual Learning with Guaranteed Model Developmental Safety
by: Li, Gang, et al.
Published: (2024)
by: Li, Gang, et al.
Published: (2024)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
by: Thekumparampil, Kiran Koshy, et al.
Published: (2024)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
by: Xu, Conglong, et al.
Published: (2025)
by: Xu, Conglong, et al.
Published: (2025)
GLinSAT: The General Linear Satisfiability Neural Network Layer By Accelerated Gradient Descent
by: Zeng, Hongtai, et al.
Published: (2024)
by: Zeng, Hongtai, et al.
Published: (2024)
Fast Computation of Optimal Transport via Entropy-Regularized Extragradient Methods
by: Li, Gen, et al.
Published: (2023)
by: Li, Gen, et al.
Published: (2023)
Online Riemannian Gradient Descent for Quantum State Tomography with Matrix Product Operators
by: Cai, Jian-Feng, et al.
Published: (2026)
by: Cai, Jian-Feng, et al.
Published: (2026)
An Improved Last-Iterate Convergence Rate for Anchored Gradient Descent Ascent
by: Surina, Anja, et al.
Published: (2026)
by: Surina, Anja, et al.
Published: (2026)
Gradient Descent Efficiency Index
by: Dhingra, Aviral
Published: (2024)
by: Dhingra, Aviral
Published: (2024)
Jacobian Descent for Multi-Objective Optimization
by: Quinton, Pierre, et al.
Published: (2024)
by: Quinton, Pierre, et al.
Published: (2024)
Capabilities and Fundamental Limits of Latent Chain-of-Thought
by: Zou, Jiaxuan, et al.
Published: (2026)
by: Zou, Jiaxuan, et al.
Published: (2026)
One if by Land, Two if by Sea, Three if by Four Seas, and More to Come -- Values of Perception, Prediction, Communication, and Common Sense in Decision Making
by: Xu, Aolin
Published: (2025)
by: Xu, Aolin
Published: (2025)
High-probability sample complexities for policy evaluation with linear function approximation
by: Li, Gen, et al.
Published: (2023)
by: Li, Gen, et al.
Published: (2023)
Sven: Singular Value Descent as a Computationally Efficient Natural Gradient Method
by: Bright-Thonney, Samuel, et al.
Published: (2026)
by: Bright-Thonney, Samuel, et al.
Published: (2026)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Queueing-Aware Optimization of Reasoning Tokens for Accuracy-Latency Trade-offs in LLM Servers
by: Ozbas, Emre, et al.
Published: (2026)
by: Ozbas, Emre, et al.
Published: (2026)
Multi-Robot Multi-Queue Control via Exhaustive Assignment Actor-Critic Learning
by: Merati, Mohammad, et al.
Published: (2026)
by: Merati, Mohammad, et al.
Published: (2026)
Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descent
by: Hennick, Max, et al.
Published: (2025)
by: Hennick, Max, et al.
Published: (2025)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
by: Gao, Yihang, et al.
Published: (2024)
by: Gao, Yihang, et al.
Published: (2024)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Similar Items
-
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026) -
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
by: Yang, Tong, et al.
Published: (2024) -
Accelerating Convergence of Score-Based Diffusion Models, Provably
by: Li, Gen, et al.
Published: (2024) -
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025) -
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
by: Yang, Tong, et al.
Published: (2025)