A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Doron-Arad, Ilan, Mehalel, Idan, Mossel, Elchanan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Online Realizable Regression and Applications for ReLU Networks
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2026)
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2026)
Why ReLU? A Bit-Model Dichotomy for Deep Network Training
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2026)
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2026)
Online Learning of Neural Networks
von: Daniely, Amit, et al.
Veröffentlicht: (2025)
von: Daniely, Amit, et al.
Veröffentlicht: (2025)
Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-End
von: Hanneke, Steve, et al.
Veröffentlicht: (2026)
von: Hanneke, Steve, et al.
Veröffentlicht: (2026)
On the Hardness of Training Deep Neural Networks Discretely
von: Doron-Arad, Ilan
Veröffentlicht: (2024)
von: Doron-Arad, Ilan
Veröffentlicht: (2024)
A Mathematical Model for Curriculum Learning for Parities
von: Cornacchia, Elisabetta, et al.
Veröffentlicht: (2023)
von: Cornacchia, Elisabetta, et al.
Veröffentlicht: (2023)
Most Convolutional Networks Suffer from Small Adversarial Perturbations
von: Daniely, Amit, et al.
Veröffentlicht: (2026)
von: Daniely, Amit, et al.
Veröffentlicht: (2026)
Deterministic Apple Tasting
von: Chase, Zachary, et al.
Veröffentlicht: (2024)
von: Chase, Zachary, et al.
Veröffentlicht: (2024)
Noise Sensitivity and Learning Lower Bounds for Hierarchical Functions
von: Li, Rupert, et al.
Veröffentlicht: (2025)
von: Li, Rupert, et al.
Veröffentlicht: (2025)
Some Theoretical Limitations of t-SNE
von: Li, Rupert, et al.
Veröffentlicht: (2026)
von: Li, Rupert, et al.
Veröffentlicht: (2026)
Bandit-Feedback Online Multiclass Classification: Variants and Tradeoffs
von: Filmus, Yuval, et al.
Veröffentlicht: (2024)
von: Filmus, Yuval, et al.
Veröffentlicht: (2024)
The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently
von: Cornacchia, Elisabetta, et al.
Veröffentlicht: (2026)
von: Cornacchia, Elisabetta, et al.
Veröffentlicht: (2026)
A Tight Lower Bound for Non-stochastic Multi-armed Bandits with Expert Advice
von: Chase, Zachary, et al.
Veröffentlicht: (2025)
von: Chase, Zachary, et al.
Veröffentlicht: (2025)
Sample-Efficient Linear Regression with Self-Selection Bias
von: Gaitonde, Jason, et al.
Veröffentlicht: (2024)
von: Gaitonde, Jason, et al.
Veröffentlicht: (2024)
Better Models and Algorithms for Learning Ising Models from Dynamics
von: Gaitonde, Jason, et al.
Veröffentlicht: (2025)
von: Gaitonde, Jason, et al.
Veröffentlicht: (2025)
Multiclass Online Learnability under Bandit Feedback
von: Raman, Ananth, et al.
Veröffentlicht: (2023)
von: Raman, Ananth, et al.
Veröffentlicht: (2023)
A Theory of Learning with Autoregressive Chain of Thought
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
Bypassing the Noisy Parity Barrier: Learning Higher-Order Markov Random Fields from Dynamics
von: Gaitonde, Jason, et al.
Veröffentlicht: (2024)
von: Gaitonde, Jason, et al.
Veröffentlicht: (2024)
Low-dimensional Functions are Efficiently Learnable under Randomly Biased Distributions
von: Cornacchia, Elisabetta, et al.
Veröffentlicht: (2025)
von: Cornacchia, Elisabetta, et al.
Veröffentlicht: (2025)
Denoising distances beyond the volumetric barrier
von: Huang, Han, et al.
Veröffentlicht: (2026)
von: Huang, Han, et al.
Veröffentlicht: (2026)
Reconstructing the Geometry of Random Geometric Graphs
von: Huang, Han, et al.
Veröffentlicht: (2024)
von: Huang, Han, et al.
Veröffentlicht: (2024)
A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning
von: Gaitonde, Jason, et al.
Veröffentlicht: (2026)
von: Gaitonde, Jason, et al.
Veröffentlicht: (2026)
Optimal Prediction Using Expert Advice and Randomized Littlestone Dimension
von: Filmus, Yuval, et al.
Veröffentlicht: (2023)
von: Filmus, Yuval, et al.
Veröffentlicht: (2023)
The Refutability Gap: Challenges in Validating Reasoning by Large Language Models
von: Mossel, Elchanan
Veröffentlicht: (2025)
von: Mossel, Elchanan
Veröffentlicht: (2025)
Learning and Testing Convex Functions
von: Pinto Jr., Renato Ferreira, et al.
Veröffentlicht: (2025)
von: Pinto Jr., Renato Ferreira, et al.
Veröffentlicht: (2025)
On Learning Verifiers and Implications to Chain-of-Thought Reasoning
von: Balcan, Maria-Florina, et al.
Veröffentlicht: (2025)
von: Balcan, Maria-Florina, et al.
Veröffentlicht: (2025)
On Algorithmic Robustness of Corrupted Markov Chains
von: Gaitonde, Jason, et al.
Veröffentlicht: (2025)
von: Gaitonde, Jason, et al.
Veröffentlicht: (2025)
Non-Linear Paging
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2024)
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2024)
Fractured Chain-of-Thought Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
von: Pengmei, Zihan, et al.
Veröffentlicht: (2025)
von: Pengmei, Zihan, et al.
Veröffentlicht: (2025)
Regret-Oracle Complexity Tradeoffs in Agnostic Online Learning
von: Attias, Idan, et al.
Veröffentlicht: (2026)
von: Attias, Idan, et al.
Veröffentlicht: (2026)
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2025)
Eliciting Chain-of-Thought Reasoning for Time Series Analysis using Reinforcement Learning
von: Parker, Felix, et al.
Veröffentlicht: (2025)
von: Parker, Felix, et al.
Veröffentlicht: (2025)
Continuous Chain of Thought Enables Parallel Exploration and Reasoning
von: Gozeten, Halil Alperen, et al.
Veröffentlicht: (2025)
von: Gozeten, Halil Alperen, et al.
Veröffentlicht: (2025)
Reasoning Models Sometimes Output Illegible Chains of Thought
von: Jose, Arun
Veröffentlicht: (2025)
von: Jose, Arun
Veröffentlicht: (2025)
RL's Razor: Why Online Reinforcement Learning Forgets Less
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
Demystifying Long Chain-of-Thought Reasoning in LLMs
von: Yeo, Edward, et al.
Veröffentlicht: (2025)
von: Yeo, Edward, et al.
Veröffentlicht: (2025)
Understanding Hidden Computations in Chain-of-Thought Reasoning
von: Bharadwaj, Aryasomayajula Ram
Veröffentlicht: (2024)
von: Bharadwaj, Aryasomayajula Ram
Veröffentlicht: (2024)
Unveiling Confirmation Bias in Chain-of-Thought Reasoning
von: Wan, Yue, et al.
Veröffentlicht: (2025)
von: Wan, Yue, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Online Realizable Regression and Applications for ReLU Networks
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2026) -
Why ReLU? A Bit-Model Dichotomy for Deep Network Training
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2026) -
Online Learning of Neural Networks
von: Daniely, Amit, et al.
Veröffentlicht: (2025) -
Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-End
von: Hanneke, Steve, et al.
Veröffentlicht: (2026) -
On the Hardness of Training Deep Neural Networks Discretely
von: Doron-Arad, Ilan
Veröffentlicht: (2024)