A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Doron-Arad, Ilan, Mehalel, Idan, Mossel, Elchanan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Online Realizable Regression and Applications for ReLU Networks
di: Doron-Arad, Ilan, et al.
Pubblicazione: (2026)
di: Doron-Arad, Ilan, et al.
Pubblicazione: (2026)
Why ReLU? A Bit-Model Dichotomy for Deep Network Training
di: Doron-Arad, Ilan, et al.
Pubblicazione: (2026)
di: Doron-Arad, Ilan, et al.
Pubblicazione: (2026)
Online Learning of Neural Networks
di: Daniely, Amit, et al.
Pubblicazione: (2025)
di: Daniely, Amit, et al.
Pubblicazione: (2025)
Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-End
di: Hanneke, Steve, et al.
Pubblicazione: (2026)
di: Hanneke, Steve, et al.
Pubblicazione: (2026)
On the Hardness of Training Deep Neural Networks Discretely
di: Doron-Arad, Ilan
Pubblicazione: (2024)
di: Doron-Arad, Ilan
Pubblicazione: (2024)
A Mathematical Model for Curriculum Learning for Parities
di: Cornacchia, Elisabetta, et al.
Pubblicazione: (2023)
di: Cornacchia, Elisabetta, et al.
Pubblicazione: (2023)
Most Convolutional Networks Suffer from Small Adversarial Perturbations
di: Daniely, Amit, et al.
Pubblicazione: (2026)
di: Daniely, Amit, et al.
Pubblicazione: (2026)
Deterministic Apple Tasting
di: Chase, Zachary, et al.
Pubblicazione: (2024)
di: Chase, Zachary, et al.
Pubblicazione: (2024)
Noise Sensitivity and Learning Lower Bounds for Hierarchical Functions
di: Li, Rupert, et al.
Pubblicazione: (2025)
di: Li, Rupert, et al.
Pubblicazione: (2025)
Some Theoretical Limitations of t-SNE
di: Li, Rupert, et al.
Pubblicazione: (2026)
di: Li, Rupert, et al.
Pubblicazione: (2026)
Bandit-Feedback Online Multiclass Classification: Variants and Tradeoffs
di: Filmus, Yuval, et al.
Pubblicazione: (2024)
di: Filmus, Yuval, et al.
Pubblicazione: (2024)
The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently
di: Cornacchia, Elisabetta, et al.
Pubblicazione: (2026)
di: Cornacchia, Elisabetta, et al.
Pubblicazione: (2026)
A Tight Lower Bound for Non-stochastic Multi-armed Bandits with Expert Advice
di: Chase, Zachary, et al.
Pubblicazione: (2025)
di: Chase, Zachary, et al.
Pubblicazione: (2025)
Sample-Efficient Linear Regression with Self-Selection Bias
di: Gaitonde, Jason, et al.
Pubblicazione: (2024)
di: Gaitonde, Jason, et al.
Pubblicazione: (2024)
Better Models and Algorithms for Learning Ising Models from Dynamics
di: Gaitonde, Jason, et al.
Pubblicazione: (2025)
di: Gaitonde, Jason, et al.
Pubblicazione: (2025)
Multiclass Online Learnability under Bandit Feedback
di: Raman, Ananth, et al.
Pubblicazione: (2023)
di: Raman, Ananth, et al.
Pubblicazione: (2023)
A Theory of Learning with Autoregressive Chain of Thought
di: Joshi, Nirmit, et al.
Pubblicazione: (2025)
di: Joshi, Nirmit, et al.
Pubblicazione: (2025)
Bypassing the Noisy Parity Barrier: Learning Higher-Order Markov Random Fields from Dynamics
di: Gaitonde, Jason, et al.
Pubblicazione: (2024)
di: Gaitonde, Jason, et al.
Pubblicazione: (2024)
Low-dimensional Functions are Efficiently Learnable under Randomly Biased Distributions
di: Cornacchia, Elisabetta, et al.
Pubblicazione: (2025)
di: Cornacchia, Elisabetta, et al.
Pubblicazione: (2025)
Denoising distances beyond the volumetric barrier
di: Huang, Han, et al.
Pubblicazione: (2026)
di: Huang, Han, et al.
Pubblicazione: (2026)
Reconstructing the Geometry of Random Geometric Graphs
di: Huang, Han, et al.
Pubblicazione: (2024)
di: Huang, Han, et al.
Pubblicazione: (2024)
A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning
di: Gaitonde, Jason, et al.
Pubblicazione: (2026)
di: Gaitonde, Jason, et al.
Pubblicazione: (2026)
Optimal Prediction Using Expert Advice and Randomized Littlestone Dimension
di: Filmus, Yuval, et al.
Pubblicazione: (2023)
di: Filmus, Yuval, et al.
Pubblicazione: (2023)
The Refutability Gap: Challenges in Validating Reasoning by Large Language Models
di: Mossel, Elchanan
Pubblicazione: (2025)
di: Mossel, Elchanan
Pubblicazione: (2025)
Learning and Testing Convex Functions
di: Pinto Jr., Renato Ferreira, et al.
Pubblicazione: (2025)
di: Pinto Jr., Renato Ferreira, et al.
Pubblicazione: (2025)
On Learning Verifiers and Implications to Chain-of-Thought Reasoning
di: Balcan, Maria-Florina, et al.
Pubblicazione: (2025)
di: Balcan, Maria-Florina, et al.
Pubblicazione: (2025)
On Algorithmic Robustness of Corrupted Markov Chains
di: Gaitonde, Jason, et al.
Pubblicazione: (2025)
di: Gaitonde, Jason, et al.
Pubblicazione: (2025)
Non-Linear Paging
di: Doron-Arad, Ilan, et al.
Pubblicazione: (2024)
di: Doron-Arad, Ilan, et al.
Pubblicazione: (2024)
Fractured Chain-of-Thought Reasoning
di: Liao, Baohao, et al.
Pubblicazione: (2025)
di: Liao, Baohao, et al.
Pubblicazione: (2025)
The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
di: Pengmei, Zihan, et al.
Pubblicazione: (2025)
di: Pengmei, Zihan, et al.
Pubblicazione: (2025)
Regret-Oracle Complexity Tradeoffs in Agnostic Online Learning
di: Attias, Idan, et al.
Pubblicazione: (2026)
di: Attias, Idan, et al.
Pubblicazione: (2026)
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
di: Zhu, Hanlin, et al.
Pubblicazione: (2025)
di: Zhu, Hanlin, et al.
Pubblicazione: (2025)
Eliciting Chain-of-Thought Reasoning for Time Series Analysis using Reinforcement Learning
di: Parker, Felix, et al.
Pubblicazione: (2025)
di: Parker, Felix, et al.
Pubblicazione: (2025)
Continuous Chain of Thought Enables Parallel Exploration and Reasoning
di: Gozeten, Halil Alperen, et al.
Pubblicazione: (2025)
di: Gozeten, Halil Alperen, et al.
Pubblicazione: (2025)
Reasoning Models Sometimes Output Illegible Chains of Thought
di: Jose, Arun
Pubblicazione: (2025)
di: Jose, Arun
Pubblicazione: (2025)
RL's Razor: Why Online Reinforcement Learning Forgets Less
di: Shenfeld, Idan, et al.
Pubblicazione: (2025)
di: Shenfeld, Idan, et al.
Pubblicazione: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
Demystifying Long Chain-of-Thought Reasoning in LLMs
di: Yeo, Edward, et al.
Pubblicazione: (2025)
di: Yeo, Edward, et al.
Pubblicazione: (2025)
Understanding Hidden Computations in Chain-of-Thought Reasoning
di: Bharadwaj, Aryasomayajula Ram
Pubblicazione: (2024)
di: Bharadwaj, Aryasomayajula Ram
Pubblicazione: (2024)
Unveiling Confirmation Bias in Chain-of-Thought Reasoning
di: Wan, Yue, et al.
Pubblicazione: (2025)
di: Wan, Yue, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Online Realizable Regression and Applications for ReLU Networks
di: Doron-Arad, Ilan, et al.
Pubblicazione: (2026) -
Why ReLU? A Bit-Model Dichotomy for Deep Network Training
di: Doron-Arad, Ilan, et al.
Pubblicazione: (2026) -
Online Learning of Neural Networks
di: Daniely, Amit, et al.
Pubblicazione: (2025) -
Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-End
di: Hanneke, Steve, et al.
Pubblicazione: (2026) -
On the Hardness of Training Deep Neural Networks Discretely
di: Doron-Arad, Ilan
Pubblicazione: (2024)