Transformers Provably Learn to Internalize Chain-of-Thought
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yixiao, Zhu, Hanlin, Wang, Zixuan, Jiao, Jiantao, Russell, Stuart, Sojoudi, Somayeh, Mei, Song |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
by: Huang, Yixiao, et al.
Published: (2025)
by: Huang, Yixiao, et al.
Published: (2025)
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
by: Zhu, Hanlin, et al.
Published: (2025)
by: Zhu, Hanlin, et al.
Published: (2025)
Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought
by: Zhu, Hanlin, et al.
Published: (2025)
by: Zhu, Hanlin, et al.
Published: (2025)
Reinforcement Learning for Flow-Matching Policies
by: Pfrommer, Samuel, et al.
Published: (2025)
by: Pfrommer, Samuel, et al.
Published: (2025)
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
by: Pfrommer, Samuel, et al.
Published: (2025)
by: Pfrommer, Samuel, et al.
Published: (2025)
Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning
by: Cheng, Ziheng, et al.
Published: (2026)
by: Cheng, Ziheng, et al.
Published: (2026)
GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
by: Zhu, Hanlin, et al.
Published: (2025)
by: Zhu, Hanlin, et al.
Published: (2025)
OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
by: Cheng, Ziheng, et al.
Published: (2025)
by: Cheng, Ziheng, et al.
Published: (2025)
On Representation Complexity of Model-based and Model-free Reinforcement Learning
by: Zhu, Hanlin, et al.
Published: (2023)
by: Zhu, Hanlin, et al.
Published: (2023)
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
by: Zhu, Hanlin, et al.
Published: (2024)
by: Zhu, Hanlin, et al.
Published: (2024)
Efficient Prompt Caching via Embedding Similarity
by: Zhu, Hanlin, et al.
Published: (2024)
by: Zhu, Hanlin, et al.
Published: (2024)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
by: Huang, Jianhao, et al.
Published: (2025)
by: Huang, Jianhao, et al.
Published: (2025)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Avoiding Catastrophe in Online Learning by Asking for Help
by: Plaut, Benjamin, et al.
Published: (2024)
by: Plaut, Benjamin, et al.
Published: (2024)
Pausing Policy Learning in Non-stationary Reinforcement Learning
by: Lee, Hyunin, et al.
Published: (2024)
by: Lee, Hyunin, et al.
Published: (2024)
Transport of Algebraic Structure to Latent Embeddings
by: Pfrommer, Samuel, et al.
Published: (2024)
by: Pfrommer, Samuel, et al.
Published: (2024)
Reinforcement Learning via Value Gradient Flow
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Safe Learning Under Irreversible Dynamics via Asking for Help
by: Plaut, Benjamin, et al.
Published: (2025)
by: Plaut, Benjamin, et al.
Published: (2025)
Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
by: Cui, Yixin, et al.
Published: (2025)
by: Cui, Yixin, et al.
Published: (2025)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
by: Ma, George, et al.
Published: (2026)
by: Ma, George, et al.
Published: (2026)
Mixing Classifiers to Alleviate the Accuracy-Robustness Trade-Off
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses
by: Ma, Ziye, et al.
Published: (2024)
by: Ma, Ziye, et al.
Published: (2024)
Towards Optimal Branching of Linear and Semidefinite Relaxations for Neural Network Robustness Certification
by: Anderson, Brendon G., et al.
Published: (2021)
by: Anderson, Brendon G., et al.
Published: (2021)
Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning
by: Li, Jingqi, et al.
Published: (2022)
by: Li, Jingqi, et al.
Published: (2022)
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
by: Huang, Sili, et al.
Published: (2024)
by: Huang, Sili, et al.
Published: (2024)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
DRAGON: Distributional Rewards Optimize Diffusion Generative Models
by: Bai, Yatong, et al.
Published: (2025)
by: Bai, Yatong, et al.
Published: (2025)
Efficient Global Optimization of Two-Layer ReLU Networks: Quadratic-Time Algorithms and Adversarial Training
by: Bai, Yatong, et al.
Published: (2022)
by: Bai, Yatong, et al.
Published: (2022)
Provably Convergent Federated Trilevel Learning
by: Jiao, Yang, et al.
Published: (2023)
by: Jiao, Yang, et al.
Published: (2023)
The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
by: Pengmei, Zihan, et al.
Published: (2025)
by: Pengmei, Zihan, et al.
Published: (2025)
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
by: Amiri, Alireza, et al.
Published: (2025)
by: Amiri, Alireza, et al.
Published: (2025)
Latent Chain-of-Thought Improves Structured-Data Transformers
by: Dudley, Carson, et al.
Published: (2026)
by: Dudley, Carson, et al.
Published: (2026)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
by: Guo, Tianyu, et al.
Published: (2024)
by: Guo, Tianyu, et al.
Published: (2024)
Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMs
by: Yang, Zhipeng, et al.
Published: (2025)
by: Yang, Zhipeng, et al.
Published: (2025)
Improving the Accuracy-Robustness Trade-Off of Classifiers via Adaptive Smoothing
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
Similar Items
-
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
by: Huang, Yixiao, et al.
Published: (2025) -
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
by: Zhu, Hanlin, et al.
Published: (2025) -
Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought
by: Zhu, Hanlin, et al.
Published: (2025) -
Reinforcement Learning for Flow-Matching Policies
by: Pfrommer, Samuel, et al.
Published: (2025) -
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
by: Pfrommer, Samuel, et al.
Published: (2025)