Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum
Fuente:
arXiv
Salvato in:
| Autori principali: | Rajaraman, Nived, Huang, Audrey, Dudik, Miro, Schapire, Robert, Foster, Dylan J., Krishnamurthy, Akshay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the Hardness of Bandit Learning
di: Brukhim, Nataly, et al.
Pubblicazione: (2025)
di: Brukhim, Nataly, et al.
Pubblicazione: (2025)
Astral Space: Convex Analysis at Infinity
di: Dudík, Miroslav, et al.
Pubblicazione: (2022)
di: Dudík, Miroslav, et al.
Pubblicazione: (2022)
Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification
di: Rohatgi, Dhruv, et al.
Pubblicazione: (2025)
di: Rohatgi, Dhruv, et al.
Pubblicazione: (2025)
Interactive Learning of Single-Index Models via Stochastic Gradient Descent
di: Rajaraman, Nived, et al.
Pubblicazione: (2026)
di: Rajaraman, Nived, et al.
Pubblicazione: (2026)
Provable Interactive Learning with Hindsight Instruction Feedback
di: Misra, Dipendra, et al.
Pubblicazione: (2024)
di: Misra, Dipendra, et al.
Pubblicazione: (2024)
Rich-Observation Reinforcement Learning with Continuous Latent Dynamics
di: Song, Yuda, et al.
Pubblicazione: (2024)
di: Song, Yuda, et al.
Pubblicazione: (2024)
Scalable Online Exploration via Coverability
di: Amortila, Philip, et al.
Pubblicazione: (2024)
di: Amortila, Philip, et al.
Pubblicazione: (2024)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
di: Bu, Dake, et al.
Pubblicazione: (2025)
di: Bu, Dake, et al.
Pubblicazione: (2025)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
di: Huang, Audrey, et al.
Pubblicazione: (2025)
di: Huang, Audrey, et al.
Pubblicazione: (2025)
Representation-Based Exploration for Language Models: From Test-Time to Post-Training
di: Tuyls, Jens, et al.
Pubblicazione: (2025)
di: Tuyls, Jens, et al.
Pubblicazione: (2025)
Toward a Theory of Tokenization in LLMs
di: Rajaraman, Nived, et al.
Pubblicazione: (2024)
di: Rajaraman, Nived, et al.
Pubblicazione: (2024)
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
di: Amortila, Philip, et al.
Pubblicazione: (2024)
di: Amortila, Philip, et al.
Pubblicazione: (2024)
The Space Complexity of Learning-Unlearning Algorithms
di: Cherapanamjeri, Yeshwanth, et al.
Pubblicazione: (2025)
di: Cherapanamjeri, Yeshwanth, et al.
Pubblicazione: (2025)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
di: Huang, Audrey, et al.
Pubblicazione: (2024)
di: Huang, Audrey, et al.
Pubblicazione: (2024)
Self-Improvement in Language Models: The Sharpening Mechanism
di: Huang, Audrey, et al.
Pubblicazione: (2024)
di: Huang, Audrey, et al.
Pubblicazione: (2024)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
di: Setlur, Amrith, et al.
Pubblicazione: (2025)
di: Setlur, Amrith, et al.
Pubblicazione: (2025)
Can large language models explore in-context?
di: Krishnamurthy, Akshay, et al.
Pubblicazione: (2024)
di: Krishnamurthy, Akshay, et al.
Pubblicazione: (2024)
A Unifying View of Coverage in Linear Off-Policy Evaluation
di: Amortila, Philip, et al.
Pubblicazione: (2026)
di: Amortila, Philip, et al.
Pubblicazione: (2026)
Computational Intractability of Strategizing against Online Learners
di: Assos, Angelos, et al.
Pubblicazione: (2025)
di: Assos, Angelos, et al.
Pubblicazione: (2025)
The Coverage Principle: How Pre-Training Enables Post-Training
di: Chen, Fan, et al.
Pubblicazione: (2025)
di: Chen, Fan, et al.
Pubblicazione: (2025)
Statistical Complexity and Optimal Algorithms for Non-linear Ridge Bandits
di: Rajaraman, Nived, et al.
Pubblicazione: (2023)
di: Rajaraman, Nived, et al.
Pubblicazione: (2023)
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
di: Ekbote, Chanakya, et al.
Pubblicazione: (2025)
di: Ekbote, Chanakya, et al.
Pubblicazione: (2025)
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
di: Golowich, Noah, et al.
Pubblicazione: (2026)
di: Golowich, Noah, et al.
Pubblicazione: (2026)
On Provable Benefits of Muon in Federated Learning
di: Zhang, Xinwen, et al.
Pubblicazione: (2025)
di: Zhang, Xinwen, et al.
Pubblicazione: (2025)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
di: Xie, Tengyang, et al.
Pubblicazione: (2024)
di: Xie, Tengyang, et al.
Pubblicazione: (2024)
Provable Benefits of Sinusoidal Activation for Modular Addition
di: Huang, Tianlong, et al.
Pubblicazione: (2025)
di: Huang, Tianlong, et al.
Pubblicazione: (2025)
Momentum Benefits Non-IID Federated Learning Simply and Provably
di: Cheng, Ziheng, et al.
Pubblicazione: (2023)
di: Cheng, Ziheng, et al.
Pubblicazione: (2023)
Transformers on Markov Data: Constant Depth Suffices
di: Rajaraman, Nived, et al.
Pubblicazione: (2024)
di: Rajaraman, Nived, et al.
Pubblicazione: (2024)
On the Benefit of Optimal Transport for Curriculum Reinforcement Learning
di: Klink, Pascal, et al.
Pubblicazione: (2023)
di: Klink, Pascal, et al.
Pubblicazione: (2023)
The Role of Environment Access in Agnostic Reinforcement Learning
di: Krishnamurthy, Akshay, et al.
Pubblicazione: (2025)
di: Krishnamurthy, Akshay, et al.
Pubblicazione: (2025)
Provable Benefits of In-Tool Learning for Large Language Models
di: Houliston, Sam, et al.
Pubblicazione: (2025)
di: Houliston, Sam, et al.
Pubblicazione: (2025)
Provable Benefit of Cutout and CutMix for Feature Learning
di: Oh, Junsoo, et al.
Pubblicazione: (2024)
di: Oh, Junsoo, et al.
Pubblicazione: (2024)
Mitigating Covariate Shift in Misspecified Regression with Applications to Reinforcement Learning
di: Amortila, Philip, et al.
Pubblicazione: (2024)
di: Amortila, Philip, et al.
Pubblicazione: (2024)
A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning
di: Gaitonde, Jason, et al.
Pubblicazione: (2026)
di: Gaitonde, Jason, et al.
Pubblicazione: (2026)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
di: Kim, Juno, et al.
Pubblicazione: (2025)
di: Kim, Juno, et al.
Pubblicazione: (2025)
Necessary and Sufficient Oracles: Toward a Computational Taxonomy For Reinforcement Learning
di: Rohatgi, Dhruv, et al.
Pubblicazione: (2025)
di: Rohatgi, Dhruv, et al.
Pubblicazione: (2025)
Wait, Wait, Wait... Why Do Reasoning Models Loop?
di: Pipis, Charilaos, et al.
Pubblicazione: (2025)
di: Pipis, Charilaos, et al.
Pubblicazione: (2025)
Provable Benefits of Unsupervised Pre-training and Transfer Learning via Single-Index Models
di: Jones-McCormick, Taj, et al.
Pubblicazione: (2025)
di: Jones-McCormick, Taj, et al.
Pubblicazione: (2025)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
di: Bondaschi, Marco, et al.
Pubblicazione: (2025)
di: Bondaschi, Marco, et al.
Pubblicazione: (2025)
The Power of Resets in Online Reinforcement Learning
di: Mhammedi, Zakaria, et al.
Pubblicazione: (2024)
di: Mhammedi, Zakaria, et al.
Pubblicazione: (2024)
Documenti analoghi
-
On the Hardness of Bandit Learning
di: Brukhim, Nataly, et al.
Pubblicazione: (2025) -
Astral Space: Convex Analysis at Infinity
di: Dudík, Miroslav, et al.
Pubblicazione: (2022) -
Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification
di: Rohatgi, Dhruv, et al.
Pubblicazione: (2025) -
Interactive Learning of Single-Index Models via Stochastic Gradient Descent
di: Rajaraman, Nived, et al.
Pubblicazione: (2026) -
Provable Interactive Learning with Hindsight Instruction Feedback
di: Misra, Dipendra, et al.
Pubblicazione: (2024)