Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Hanlin, Hao, Shibo, Hu, Zhiting, Jiao, Jiantao, Russell, Stuart, Tian, Yuandong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
Transformers Provably Learn to Internalize Chain-of-Thought
von: Huang, Yixiao, et al.
Veröffentlicht: (2026)
von: Huang, Yixiao, et al.
Veröffentlicht: (2026)
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
von: Tian, Yuandong
Veröffentlicht: (2025)
von: Tian, Yuandong
Veröffentlicht: (2025)
Efficient Prompt Caching via Embedding Similarity
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
von: Su, DiJia, et al.
Veröffentlicht: (2025)
von: Su, DiJia, et al.
Veröffentlicht: (2025)
GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
On Representation Complexity of Model-based and Model-free Reinforcement Learning
von: Zhu, Hanlin, et al.
Veröffentlicht: (2023)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2023)
ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
von: Hao, Shibo, et al.
Veröffentlicht: (2023)
von: Hao, Shibo, et al.
Veröffentlicht: (2023)
Avoiding Catastrophe in Online Learning by Asking for Help
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
von: Plaut, Benjamin, et al.
Veröffentlicht: (2024)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
Safe Learning Under Irreversible Dynamics via Asking for Help
von: Plaut, Benjamin, et al.
Veröffentlicht: (2025)
von: Plaut, Benjamin, et al.
Veröffentlicht: (2025)
Deep Thinking by Markov Chain of Continuous Thoughts
von: Liu, Jiayu, et al.
Veröffentlicht: (2025)
von: Liu, Jiayu, et al.
Veröffentlicht: (2025)
Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought
von: Xie, Zixuan, et al.
Veröffentlicht: (2026)
von: Xie, Zixuan, et al.
Veröffentlicht: (2026)
Unveiling Confirmation Bias in Chain-of-Thought Reasoning
von: Wan, Yue, et al.
Veröffentlicht: (2025)
von: Wan, Yue, et al.
Veröffentlicht: (2025)
Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
von: Hu, Xinyang, et al.
Veröffentlicht: (2024)
von: Hu, Xinyang, et al.
Veröffentlicht: (2024)
LLM Pretraining with Continuous Concepts
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
von: Tack, Jihoon, et al.
Veröffentlicht: (2025)
Training Large Language Models to Reason in a Continuous Latent Space
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
Continuous Chain of Thought Enables Parallel Exploration and Reasoning
von: Gozeten, Halil Alperen, et al.
Veröffentlicht: (2025)
von: Gozeten, Halil Alperen, et al.
Veröffentlicht: (2025)
Emergence of Frontier Superposition: Möbius attractor and Cascade Supervision
von: Gu, Hongyu, et al.
Veröffentlicht: (2026)
von: Gu, Hongyu, et al.
Veröffentlicht: (2026)
Constraint-Rectified Training for Efficient Chain-of-Thought
von: Wu, Qinhang, et al.
Veröffentlicht: (2026)
von: Wu, Qinhang, et al.
Veröffentlicht: (2026)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
Understanding the Effect of Noise in LLM Training Data with Algorithmic Chains of Thought
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
Composing Global Solutions to Reasoning Tasks via Algebraic Objects in Neural Nets
von: Tian, Yuandong
Veröffentlicht: (2024)
von: Tian, Yuandong
Veröffentlicht: (2024)
Chain-of-Thought Predictive Control
von: Jia, Zhiwei, et al.
Veröffentlicht: (2023)
von: Jia, Zhiwei, et al.
Veröffentlicht: (2023)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
Dynamic Chain-of-Thought: Towards Adaptive Deep Reasoning
von: Wang, Libo
Veröffentlicht: (2025)
von: Wang, Libo
Veröffentlicht: (2025)
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
von: Zhao, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhao, Jiawei, et al.
Veröffentlicht: (2024)
Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression
von: Tang, Yuntian, et al.
Veröffentlicht: (2026)
von: Tang, Yuntian, et al.
Veröffentlicht: (2026)
Towards Optimal Statistical Watermarking
von: Huang, Baihe, et al.
Veröffentlicht: (2023)
von: Huang, Baihe, et al.
Veröffentlicht: (2023)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
von: Jia, Jinghan, et al.
Veröffentlicht: (2026)
von: Jia, Jinghan, et al.
Veröffentlicht: (2026)
Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
von: Yang, Tianyun, et al.
Veröffentlicht: (2025)
von: Yang, Tianyun, et al.
Veröffentlicht: (2025)
Toward a Theory of Tokenization in LLMs
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
dUltra: Ultra-Fast Diffusion Language Models via Reinforcement Learning
von: Chen, Shirui, et al.
Veröffentlicht: (2025)
von: Chen, Shirui, et al.
Veröffentlicht: (2025)
RedCoast: A Lightweight Tool to Automate Distributed Training of LLMs on Any GPU/TPUs
von: Tan, Bowen, et al.
Veröffentlicht: (2023)
von: Tan, Bowen, et al.
Veröffentlicht: (2023)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
Superpositional Gradient Descent: Harnessing Quantum Principles for Model Training
von: Pamuk, Ahmet Erdem, et al.
Veröffentlicht: (2025)
von: Pamuk, Ahmet Erdem, et al.
Veröffentlicht: (2025)
Inference-Time Chain-of-Thought Pruning with Latent Informativeness Signals
von: Li, Sophie, et al.
Veröffentlicht: (2025)
von: Li, Sophie, et al.
Veröffentlicht: (2025)
Robust Fully-Asynchronous Methods for Distributed Training over General Architecture
von: Zhu, Zehan, et al.
Veröffentlicht: (2023)
von: Zhu, Zehan, et al.
Veröffentlicht: (2023)
Beyond the Prompt in Large Language Models: Comprehension, In-Context Learning, and Chain-of-Thought
von: Jiao, Yuling, et al.
Veröffentlicht: (2026)
von: Jiao, Yuling, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025) -
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024) -
Transformers Provably Learn to Internalize Chain-of-Thought
von: Huang, Yixiao, et al.
Veröffentlicht: (2026) -
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
von: Tian, Yuandong
Veröffentlicht: (2025) -
Efficient Prompt Caching via Embedding Similarity
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)