Not only where, But when: Temporal Scheduling for RLVR
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Jinghao, Li, Ruilin, Zhao, Feng, Wang, Jiaqi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
por: Zhang, Jinghao, et al.
Publicado: (2025)
por: Zhang, Jinghao, et al.
Publicado: (2025)
Self-Distilled RLVR
por: Yang, Chenxu, et al.
Publicado: (2026)
por: Yang, Chenxu, et al.
Publicado: (2026)
Evaluating Parameter Efficient Methods for RLVR
por: Yin, Qingyu, et al.
Publicado: (2025)
por: Yin, Qingyu, et al.
Publicado: (2025)
Achievable distributional robustness when the robust risk is only partially identified
por: Kostin, Julia, et al.
Publicado: (2025)
por: Kostin, Julia, et al.
Publicado: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
por: Duo, Jiangshan, et al.
Publicado: (2026)
por: Duo, Jiangshan, et al.
Publicado: (2026)
The Unlearnability Phenomenon in RLVR for Language Models
por: Chen, Yulin, et al.
Publicado: (2026)
por: Chen, Yulin, et al.
Publicado: (2026)
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
por: Min, Zijun, et al.
Publicado: (2026)
por: Min, Zijun, et al.
Publicado: (2026)
RLVR-World: Training World Models with Reinforcement Learning
por: Wu, Jialong, et al.
Publicado: (2025)
por: Wu, Jialong, et al.
Publicado: (2025)
Data-Efficient RLVR via Off-Policy Influence Guidance
por: Zhu, Erle, et al.
Publicado: (2025)
por: Zhu, Erle, et al.
Publicado: (2025)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
por: Lochab, Anamika, et al.
Publicado: (2026)
por: Lochab, Anamika, et al.
Publicado: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
por: Wu, Junkang, et al.
Publicado: (2025)
por: Wu, Junkang, et al.
Publicado: (2025)
The Path Not Taken: RLVR Provably Learns Off the Principals
por: Zhu, Hanqing, et al.
Publicado: (2025)
por: Zhu, Hanqing, et al.
Publicado: (2025)
Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
por: Zhang, Yuheng, et al.
Publicado: (2025)
por: Zhang, Yuheng, et al.
Publicado: (2025)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
por: Ye, Hao, et al.
Publicado: (2026)
por: Ye, Hao, et al.
Publicado: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
por: He, Bingxiang, et al.
Publicado: (2026)
por: He, Bingxiang, et al.
Publicado: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
por: He, Zhiyuan, et al.
Publicado: (2025)
por: He, Zhiyuan, et al.
Publicado: (2025)
Dual Consensus: Escaping from Spurious Majority in Unsupervised RLVR via Two-Stage Vote Mechanism
por: Du, Kaixuan, et al.
Publicado: (2026)
por: Du, Kaixuan, et al.
Publicado: (2026)
Spurious Rewards: Rethinking Training Signals in RLVR
por: Shao, Rulin, et al.
Publicado: (2025)
por: Shao, Rulin, et al.
Publicado: (2025)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
por: Wang, Tao, et al.
Publicado: (2026)
por: Wang, Tao, et al.
Publicado: (2026)
GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR
por: Zhang, Jiaying, et al.
Publicado: (2026)
por: Zhang, Jiaying, et al.
Publicado: (2026)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
por: Lu, Han, et al.
Publicado: (2025)
por: Lu, Han, et al.
Publicado: (2025)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
por: Huang, Kexin, et al.
Publicado: (2026)
por: Huang, Kexin, et al.
Publicado: (2026)
Efficient RLVR Training via Weighted Mutual Information Data Selection
por: Zhou, Xinyu, et al.
Publicado: (2026)
por: Zhou, Xinyu, et al.
Publicado: (2026)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
por: Zhai, Zepeng, et al.
Publicado: (2026)
por: Zhai, Zepeng, et al.
Publicado: (2026)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
por: Hao, Zhezheng, et al.
Publicado: (2025)
por: Hao, Zhezheng, et al.
Publicado: (2025)
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
por: Zhang, Yiqi, et al.
Publicado: (2026)
por: Zhang, Yiqi, et al.
Publicado: (2026)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
por: Xu, Huimin, et al.
Publicado: (2026)
por: Xu, Huimin, et al.
Publicado: (2026)
Differentiable Combinatorial Scheduling at Scale
por: Liu, Mingju, et al.
Publicado: (2024)
por: Liu, Mingju, et al.
Publicado: (2024)
Linear Dynamics in the RLVR Training of Large Language Models
por: Wang, Tianle, et al.
Publicado: (2026)
por: Wang, Tianle, et al.
Publicado: (2026)
Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
por: Qiu, Zhongxi, et al.
Publicado: (2025)
por: Qiu, Zhongxi, et al.
Publicado: (2025)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
por: Zhao, Jiale, et al.
Publicado: (2026)
por: Zhao, Jiale, et al.
Publicado: (2026)
Method of data forward generation with partial differential equations for machine learning modeling in fluid mechanics
por: Chen, Ruilin
Publicado: (2025)
por: Chen, Ruilin
Publicado: (2025)
Augmenting Offline Reinforcement Learning with State-only Interactions
por: Li, Shangzhe, et al.
Publicado: (2024)
por: Li, Shangzhe, et al.
Publicado: (2024)
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
por: Mou, Chaoli, et al.
Publicado: (2026)
por: Mou, Chaoli, et al.
Publicado: (2026)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
por: Miao, Yuchun, et al.
Publicado: (2026)
por: Miao, Yuchun, et al.
Publicado: (2026)
Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection
por: Wu, Jianghao, et al.
Publicado: (2026)
por: Wu, Jianghao, et al.
Publicado: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
por: Huang, Fanding, et al.
Publicado: (2025)
por: Huang, Fanding, et al.
Publicado: (2025)
Limits of Generalization in RLVR: Two Case Studies in Mathematical Reasoning
por: Alam, Md Tanvirul, et al.
Publicado: (2025)
por: Alam, Md Tanvirul, et al.
Publicado: (2025)
Generalization of RLVR Using Causal Reasoning as a Testbed
por: Lu, Brian, et al.
Publicado: (2025)
por: Lu, Brian, et al.
Publicado: (2025)
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
por: Li, Yuhan, et al.
Publicado: (2026)
por: Li, Yuhan, et al.
Publicado: (2026)
Ejemplares similares
-
RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
por: Zhang, Jinghao, et al.
Publicado: (2025) -
Self-Distilled RLVR
por: Yang, Chenxu, et al.
Publicado: (2026) -
Evaluating Parameter Efficient Methods for RLVR
por: Yin, Qingyu, et al.
Publicado: (2025) -
Achievable distributional robustness when the robust risk is only partially identified
por: Kostin, Julia, et al.
Publicado: (2025) -
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
por: Duo, Jiangshan, et al.
Publicado: (2026)