Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Jianghao, Cai, Jianfei, Wang, Weiqiang, Ye, Jin, Schmidt, Daniel F., George, Yasmeen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
por: Wu, Jianghao, et al.
Publicado: (2025)
por: Wu, Jianghao, et al.
Publicado: (2025)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
por: Wang, Tao, et al.
Publicado: (2026)
por: Wang, Tao, et al.
Publicado: (2026)
Linear Dynamics in the RLVR Training of Large Language Models
por: Wang, Tianle, et al.
Publicado: (2026)
por: Wang, Tianle, et al.
Publicado: (2026)
Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
por: Zhang, Yuheng, et al.
Publicado: (2025)
por: Zhang, Yuheng, et al.
Publicado: (2025)
Efficient RLVR Training via Weighted Mutual Information Data Selection
por: Zhou, Xinyu, et al.
Publicado: (2026)
por: Zhou, Xinyu, et al.
Publicado: (2026)
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
por: Kim, Soeun, et al.
Publicado: (2026)
por: Kim, Soeun, et al.
Publicado: (2026)
RLVR-World: Training World Models with Reinforcement Learning
por: Wu, Jialong, et al.
Publicado: (2025)
por: Wu, Jialong, et al.
Publicado: (2025)
Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs
por: Khosravi, Hamed, et al.
Publicado: (2026)
por: Khosravi, Hamed, et al.
Publicado: (2026)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
por: Gai, Jiading, et al.
Publicado: (2026)
por: Gai, Jiading, et al.
Publicado: (2026)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
por: Ye, Hao, et al.
Publicado: (2026)
por: Ye, Hao, et al.
Publicado: (2026)
Spurious Rewards: Rethinking Training Signals in RLVR
por: Shao, Rulin, et al.
Publicado: (2025)
por: Shao, Rulin, et al.
Publicado: (2025)
Scaling Adversarial Training via Data Selection
por: Ye, Youran, et al.
Publicado: (2025)
por: Ye, Youran, et al.
Publicado: (2025)
BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning
por: Gong, Shijin, et al.
Publicado: (2026)
por: Gong, Shijin, et al.
Publicado: (2026)
Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
Learning from the Right Rollouts: Data Attribution for PPO-based LLM Post-Training
por: Shu, Dong, et al.
Publicado: (2026)
por: Shu, Dong, et al.
Publicado: (2026)
RL in the Wild: Characterizing RLVR Training in LLM Deployment
por: Zhou, Jiecheng, et al.
Publicado: (2025)
por: Zhou, Jiecheng, et al.
Publicado: (2025)
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
por: Min, Zijun, et al.
Publicado: (2026)
por: Min, Zijun, et al.
Publicado: (2026)
Data-Efficient RLVR via Off-Policy Influence Guidance
por: Zhu, Erle, et al.
Publicado: (2025)
por: Zhu, Erle, et al.
Publicado: (2025)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
por: Lu, Han, et al.
Publicado: (2025)
por: Lu, Han, et al.
Publicado: (2025)
Heddle: A Distributed Orchestration System for Agentic RL Rollout
por: Zhang, Zili, et al.
Publicado: (2026)
por: Zhang, Zili, et al.
Publicado: (2026)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
por: Miao, Yuchun, et al.
Publicado: (2026)
por: Miao, Yuchun, et al.
Publicado: (2026)
Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
por: Qiu, Zhongxi, et al.
Publicado: (2025)
por: Qiu, Zhongxi, et al.
Publicado: (2025)
Evaluating Parameter Efficient Methods for RLVR
por: Yin, Qingyu, et al.
Publicado: (2025)
por: Yin, Qingyu, et al.
Publicado: (2025)
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
por: Li, Yuhan, et al.
Publicado: (2026)
por: Li, Yuhan, et al.
Publicado: (2026)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
por: Du, Zilin, et al.
Publicado: (2024)
por: Du, Zilin, et al.
Publicado: (2024)
Self-Distilled RLVR
por: Yang, Chenxu, et al.
Publicado: (2026)
por: Yang, Chenxu, et al.
Publicado: (2026)
Efficient Data Selection for Training Genomic Perturbation Models
por: Panagopoulos, George, et al.
Publicado: (2025)
por: Panagopoulos, George, et al.
Publicado: (2025)
How Far Can Unsupervised RLVR Scale LLM Training?
por: He, Bingxiang, et al.
Publicado: (2026)
por: He, Bingxiang, et al.
Publicado: (2026)
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
por: Mou, Chaoli, et al.
Publicado: (2026)
por: Mou, Chaoli, et al.
Publicado: (2026)
Gradient-Informed Temporal Sampling Improves Rollout Accuracy in PDE Surrogate Training
por: Wang, Wenshuo, et al.
Publicado: (2026)
por: Wang, Wenshuo, et al.
Publicado: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
por: Wu, Junkang, et al.
Publicado: (2025)
por: Wu, Junkang, et al.
Publicado: (2025)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
por: Gu, Hengrui, et al.
Publicado: (2026)
por: Gu, Hengrui, et al.
Publicado: (2026)
Rollout-Training Co-Design for Efficient LLM-Based Multi-Agent Reinforcement Learning
por: Jiang, Zhida, et al.
Publicado: (2026)
por: Jiang, Zhida, et al.
Publicado: (2026)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
por: Xiong, Zidi, et al.
Publicado: (2026)
por: Xiong, Zidi, et al.
Publicado: (2026)
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
por: Sun, Yifan, et al.
Publicado: (2025)
por: Sun, Yifan, et al.
Publicado: (2025)
Not only where, But when: Temporal Scheduling for RLVR
por: Zhang, Jinghao, et al.
Publicado: (2026)
por: Zhang, Jinghao, et al.
Publicado: (2026)
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR
por: Nayak, Anupam, et al.
Publicado: (2026)
por: Nayak, Anupam, et al.
Publicado: (2026)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
por: Huang, Kexin, et al.
Publicado: (2026)
por: Huang, Kexin, et al.
Publicado: (2026)
Unraveling the Hidden Dynamical Structure in Recurrent Neural Policies
por: Li, Jin, et al.
Publicado: (2026)
por: Li, Jin, et al.
Publicado: (2026)
Intelligent Elastic Feature Fading: Enabling Model Retrain-Free Feature Efficiency Rollouts at Scale
por: Di, Jieming, et al.
Publicado: (2026)
por: Di, Jieming, et al.
Publicado: (2026)
Ejemplares similares
-
SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
por: Wu, Jianghao, et al.
Publicado: (2025) -
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
por: Wang, Tao, et al.
Publicado: (2026) -
Linear Dynamics in the RLVR Training of Large Language Models
por: Wang, Tianle, et al.
Publicado: (2026) -
Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
por: Zhang, Yuheng, et al.
Publicado: (2025) -
Efficient RLVR Training via Weighted Mutual Information Data Selection
por: Zhou, Xinyu, et al.
Publicado: (2026)