Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Hongkang, Lu, Songtao, Chen, Pin-Yu, Cui, Xiaodong, Wang, Meng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
di: Li, Hongkang, et al.
Pubblicazione: (2024)
di: Li, Hongkang, et al.
Pubblicazione: (2024)
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
di: Li, Hongkang, et al.
Pubblicazione: (2025)
di: Li, Hongkang, et al.
Pubblicazione: (2025)
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
di: Saif, A F M, et al.
Pubblicazione: (2024)
di: Saif, A F M, et al.
Pubblicazione: (2024)
When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
di: Li, Hongkang, et al.
Pubblicazione: (2025)
di: Li, Hongkang, et al.
Pubblicazione: (2025)
A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware Demonstration
di: Cui, Yingqian, et al.
Pubblicazione: (2024)
di: Cui, Yingqian, et al.
Pubblicazione: (2024)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
di: Lu, Wenquan, et al.
Pubblicazione: (2025)
di: Lu, Wenquan, et al.
Pubblicazione: (2025)
Constraint-Rectified Training for Efficient Chain-of-Thought
di: Wu, Qinhang, et al.
Pubblicazione: (2026)
di: Wu, Qinhang, et al.
Pubblicazione: (2026)
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
di: Li, Hongkang, et al.
Pubblicazione: (2024)
di: Li, Hongkang, et al.
Pubblicazione: (2024)
C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness
di: Kang, Yu, et al.
Pubblicazione: (2024)
di: Kang, Yu, et al.
Pubblicazione: (2024)
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
di: Wang, Xinyuan, et al.
Pubblicazione: (2026)
di: Wang, Xinyuan, et al.
Pubblicazione: (2026)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
di: Jia, Jinghan, et al.
Pubblicazione: (2026)
di: Jia, Jinghan, et al.
Pubblicazione: (2026)
The Expressive Power of Transformers with Chain of Thought
di: Merrill, William, et al.
Pubblicazione: (2023)
di: Merrill, William, et al.
Pubblicazione: (2023)
Enhancing Generalization in Chain of Thought Reasoning for Smaller Models
di: Yin, Maxwell J., et al.
Pubblicazione: (2025)
di: Yin, Maxwell J., et al.
Pubblicazione: (2025)
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
di: Li, Hongkang, et al.
Pubblicazione: (2024)
di: Li, Hongkang, et al.
Pubblicazione: (2024)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
di: Yehudai, Gilad, et al.
Pubblicazione: (2025)
di: Yehudai, Gilad, et al.
Pubblicazione: (2025)
Dissecting Long-Chain-of-Thought Reasoning Models: An Empirical Study
di: Mu, Yongyu, et al.
Pubblicazione: (2025)
di: Mu, Yongyu, et al.
Pubblicazione: (2025)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
di: Mohtashami, Amirkeivan, et al.
Pubblicazione: (2023)
di: Mohtashami, Amirkeivan, et al.
Pubblicazione: (2023)
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
di: Zhang, Zeliang, et al.
Pubblicazione: (2026)
di: Zhang, Zeliang, et al.
Pubblicazione: (2026)
Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning
di: Gu, Yu, et al.
Pubblicazione: (2026)
di: Gu, Yu, et al.
Pubblicazione: (2026)
Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thought
di: Chen, Zhikang, et al.
Pubblicazione: (2025)
di: Chen, Zhikang, et al.
Pubblicazione: (2025)
Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs
di: Jin, Bowen, et al.
Pubblicazione: (2024)
di: Jin, Bowen, et al.
Pubblicazione: (2024)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
di: Brösamle, Moritz, et al.
Pubblicazione: (2026)
di: Brösamle, Moritz, et al.
Pubblicazione: (2026)
Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State Tracking
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
di: Zhang, Zhen-Yu, et al.
Pubblicazione: (2024)
di: Zhang, Zhen-Yu, et al.
Pubblicazione: (2024)
Leveraging Large Language Models with Chain-of-Thought and Prompt Engineering for Traffic Crash Severity Analysis and Inference
di: Zhen, Hao, et al.
Pubblicazione: (2024)
di: Zhen, Hao, et al.
Pubblicazione: (2024)
Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling
di: Xu, Weijia, et al.
Pubblicazione: (2023)
di: Xu, Weijia, et al.
Pubblicazione: (2023)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
di: Ye, Jiacheng, et al.
Pubblicazione: (2024)
di: Ye, Jiacheng, et al.
Pubblicazione: (2024)
Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints
di: Cui, Xiaodong, et al.
Pubblicazione: (2025)
di: Cui, Xiaodong, et al.
Pubblicazione: (2025)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
di: Li, Zihao, et al.
Pubblicazione: (2025)
di: Li, Zihao, et al.
Pubblicazione: (2025)
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models
di: Shandirasegaran, Mugunthan, et al.
Pubblicazione: (2026)
di: Shandirasegaran, Mugunthan, et al.
Pubblicazione: (2026)
Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning
di: Li, Xintong, et al.
Pubblicazione: (2026)
di: Li, Xintong, et al.
Pubblicazione: (2026)
SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning
di: Hu, Chenzhi, et al.
Pubblicazione: (2026)
di: Hu, Chenzhi, et al.
Pubblicazione: (2026)
Fractured Chain-of-Thought Reasoning
di: Liao, Baohao, et al.
Pubblicazione: (2025)
di: Liao, Baohao, et al.
Pubblicazione: (2025)
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
di: Wu, Zijian, et al.
Pubblicazione: (2025)
di: Wu, Zijian, et al.
Pubblicazione: (2025)
Reinforcement Learning for Chain of Thought Compression with One-Domain-to-All Generalization
di: Li, Hanyu, et al.
Pubblicazione: (2025)
di: Li, Hanyu, et al.
Pubblicazione: (2025)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
Demystifying Long Chain-of-Thought Reasoning in LLMs
di: Yeo, Edward, et al.
Pubblicazione: (2025)
di: Yeo, Edward, et al.
Pubblicazione: (2025)
Understanding Hidden Computations in Chain-of-Thought Reasoning
di: Bharadwaj, Aryasomayajula Ram
Pubblicazione: (2024)
di: Bharadwaj, Aryasomayajula Ram
Pubblicazione: (2024)
Reliable Chain-of-Thought via Prefix Consistency
di: Iwase, Naoto, et al.
Pubblicazione: (2026)
di: Iwase, Naoto, et al.
Pubblicazione: (2026)
Documenti analoghi
-
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
di: Li, Hongkang, et al.
Pubblicazione: (2024) -
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
di: Li, Hongkang, et al.
Pubblicazione: (2025) -
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
di: Saif, A F M, et al.
Pubblicazione: (2024) -
When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
di: Li, Hongkang, et al.
Pubblicazione: (2025) -
A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware Demonstration
di: Cui, Yingqian, et al.
Pubblicazione: (2024)