The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Yu, Wen, Zixin, Chi, Yuejie, Wei, Yuting, Singh, Aarti, Liang, Yingbin, Chen, Yuxin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
di: Huang, Yu, et al.
Pubblicazione: (2025)
di: Huang, Yu, et al.
Pubblicazione: (2025)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
di: Yang, Tong, et al.
Pubblicazione: (2026)
di: Yang, Tong, et al.
Pubblicazione: (2026)
Statistical and Algorithmic Foundations of Reinforcement Learning
di: Chi, Yuejie, et al.
Pubblicazione: (2025)
di: Chi, Yuejie, et al.
Pubblicazione: (2025)
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
di: Huang, Yu, et al.
Pubblicazione: (2024)
di: Huang, Yu, et al.
Pubblicazione: (2024)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
di: Yang, Tong, et al.
Pubblicazione: (2025)
di: Yang, Tong, et al.
Pubblicazione: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
di: Ma, Jianhao, et al.
Pubblicazione: (2026)
di: Ma, Jianhao, et al.
Pubblicazione: (2026)
Accelerating Convergence of Score-Based Diffusion Models, Provably
di: Li, Gen, et al.
Pubblicazione: (2024)
di: Li, Gen, et al.
Pubblicazione: (2024)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
di: Yang, Tong, et al.
Pubblicazione: (2024)
di: Yang, Tong, et al.
Pubblicazione: (2024)
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
di: Yang, Tong, et al.
Pubblicazione: (2025)
di: Yang, Tong, et al.
Pubblicazione: (2025)
A Retention-Centric Framework for Continual Learning with Guaranteed Model Developmental Safety
di: Li, Gang, et al.
Pubblicazione: (2024)
di: Li, Gang, et al.
Pubblicazione: (2024)
Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
di: Li, Gen, et al.
Pubblicazione: (2020)
di: Li, Gen, et al.
Pubblicazione: (2020)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
di: Cen, Shicong, et al.
Pubblicazione: (2024)
di: Cen, Shicong, et al.
Pubblicazione: (2024)
Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis
di: Li, Gen, et al.
Pubblicazione: (2021)
di: Li, Gen, et al.
Pubblicazione: (2021)
Quadrupedal Robot Skateboard Mounting via Reverse Curriculum Learning
di: Belov, Danil, et al.
Pubblicazione: (2025)
di: Belov, Danil, et al.
Pubblicazione: (2025)
Deep Reinforcement Learning Optimization for Uncertain Nonlinear Systems via Event-Triggered Robust Adaptive Dynamic Programming
di: Bai, Ningwei, et al.
Pubblicazione: (2025)
di: Bai, Ningwei, et al.
Pubblicazione: (2025)
Reward-Directed Score-Based Diffusion Models via q-Learning
di: Gao, Xuefeng, et al.
Pubblicazione: (2024)
di: Gao, Xuefeng, et al.
Pubblicazione: (2024)
ODE-based Learning to Optimize
di: Xie, Zhonglin, et al.
Pubblicazione: (2024)
di: Xie, Zhonglin, et al.
Pubblicazione: (2024)
Learning to Cut: Reinforcement Learning for Benders Decomposition
di: Cai, Haochen, et al.
Pubblicazione: (2026)
di: Cai, Haochen, et al.
Pubblicazione: (2026)
Robust Evolutionary Multi-Objective Network Architecture Search for Reinforcement Learning (EMNAS-RL)
di: Adde, Nihal Acharya, et al.
Pubblicazione: (2025)
di: Adde, Nihal Acharya, et al.
Pubblicazione: (2025)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
di: Sheen, Heejune, et al.
Pubblicazione: (2024)
di: Sheen, Heejune, et al.
Pubblicazione: (2024)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
di: Rashidinejad, Paria, et al.
Pubblicazione: (2024)
di: Rashidinejad, Paria, et al.
Pubblicazione: (2024)
The Sample-Communication Complexity Trade-off in Federated Q-Learning
di: Salgia, Sudeep, et al.
Pubblicazione: (2024)
di: Salgia, Sudeep, et al.
Pubblicazione: (2024)
Aerial Inspection Behaviors via RL-based Quadrotor Control for Under-canopy Forest Environments
di: Suarez, Fausto Mauricio Lagos, et al.
Pubblicazione: (2026)
di: Suarez, Fausto Mauricio Lagos, et al.
Pubblicazione: (2026)
Model Predictive Control and Reinforcement Learning: A Unified Framework Based on Dynamic Programming
di: Bertsekas, Dimitri P.
Pubblicazione: (2024)
di: Bertsekas, Dimitri P.
Pubblicazione: (2024)
On the Implicit Bias of Adam
di: Cattaneo, Matias D., et al.
Pubblicazione: (2023)
di: Cattaneo, Matias D., et al.
Pubblicazione: (2023)
Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial
di: Li, Haoyu, et al.
Pubblicazione: (2026)
di: Li, Haoyu, et al.
Pubblicazione: (2026)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
di: Cattaneo, Matias D., et al.
Pubblicazione: (2025)
di: Cattaneo, Matias D., et al.
Pubblicazione: (2025)
Hereditary Geometric Meta-RL: Nonlocal Generalization via Task Symmetries
di: Nitschke, Paul, et al.
Pubblicazione: (2026)
di: Nitschke, Paul, et al.
Pubblicazione: (2026)
Accelerating RLHF Training with Reward Variance Increase
di: Yang, Zonglin, et al.
Pubblicazione: (2025)
di: Yang, Zonglin, et al.
Pubblicazione: (2025)
ResearchEVO: An End-to-End Framework for Automated Scientific Discovery and Documentation
di: Zhao, Zhe, et al.
Pubblicazione: (2026)
di: Zhao, Zhe, et al.
Pubblicazione: (2026)
A Beam Search Based Parallel Algorithm for the Two-Dimensional Strip Packing Problem
di: Wen, Yajie, et al.
Pubblicazione: (2025)
di: Wen, Yajie, et al.
Pubblicazione: (2025)
A Block-Based Heuristic Algorithm for the Three-Dimensional Nuclear Waste Packing Problem
di: Wen, Yajie, et al.
Pubblicazione: (2025)
di: Wen, Yajie, et al.
Pubblicazione: (2025)
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
di: Yang, Tong, et al.
Pubblicazione: (2025)
di: Yang, Tong, et al.
Pubblicazione: (2025)
Agent-Based Decentralized Energy Management of EV Charging Station with Solar Photovoltaics via Multi-Agent Reinforcement Learning
di: Fan, Jiarong, et al.
Pubblicazione: (2025)
di: Fan, Jiarong, et al.
Pubblicazione: (2025)
Deep Reinforcement Learning for Flexible Job Shop Scheduling with Random Job Arrivals
di: Tang, Yu, et al.
Pubblicazione: (2026)
di: Tang, Yu, et al.
Pubblicazione: (2026)
Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems
di: Huang, Yilie, et al.
Pubblicazione: (2025)
di: Huang, Yilie, et al.
Pubblicazione: (2025)
Optimal Transportation by Orthogonal Coupling Dynamics
di: Sadr, Mohsen, et al.
Pubblicazione: (2024)
di: Sadr, Mohsen, et al.
Pubblicazione: (2024)
Safe Decentralized Operation of EV Virtual Power Plant with Limited Network Visibility via Multi-Agent Reinforcement Learning
di: Huang, Chenghao, et al.
Pubblicazione: (2026)
di: Huang, Chenghao, et al.
Pubblicazione: (2026)
A Method to Improve the Performance of Reinforcement Learning Based on the Y Operator for a Class of Stochastic Differential Equation-Based Child-Mother Systems
di: Yin, Cheng, et al.
Pubblicazione: (2023)
di: Yin, Cheng, et al.
Pubblicazione: (2023)
Temporal Robustness in Discrete Time Linear Dynamical Systems
di: Metya, Nilava, et al.
Pubblicazione: (2025)
di: Metya, Nilava, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
di: Huang, Yu, et al.
Pubblicazione: (2025) -
Agentic Transformers Provably Learn to Search via Reinforcement Learning
di: Yang, Tong, et al.
Pubblicazione: (2026) -
Statistical and Algorithmic Foundations of Reinforcement Learning
di: Chi, Yuejie, et al.
Pubblicazione: (2025) -
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
di: Huang, Yu, et al.
Pubblicazione: (2024) -
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
di: Yang, Tong, et al.
Pubblicazione: (2025)