Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
Fuente:
arXiv
Guardado en:
| Autores principales: | Qiu, Zhongxi, Zhang, Zhang, Hu, Yan, Li, Heng, Liu, Jiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
por: Liu, Zhanyu, et al.
Publicado: (2026)
por: Liu, Zhanyu, et al.
Publicado: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
por: He, Zhiyuan, et al.
Publicado: (2025)
por: He, Zhiyuan, et al.
Publicado: (2025)
Choosing How to Remember: Adaptive Memory Structures for LLM Agents
por: Lu, Mingfei, et al.
Publicado: (2026)
por: Lu, Mingfei, et al.
Publicado: (2026)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
por: Lu, Han, et al.
Publicado: (2025)
por: Lu, Han, et al.
Publicado: (2025)
Spurious Rewards: Rethinking Training Signals in RLVR
por: Shao, Rulin, et al.
Publicado: (2025)
por: Shao, Rulin, et al.
Publicado: (2025)
RL in the Wild: Characterizing RLVR Training in LLM Deployment
por: Zhou, Jiecheng, et al.
Publicado: (2025)
por: Zhou, Jiecheng, et al.
Publicado: (2025)
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
por: Li, Yuhan, et al.
Publicado: (2026)
por: Li, Yuhan, et al.
Publicado: (2026)
RLVR-World: Training World Models with Reinforcement Learning
por: Wu, Jialong, et al.
Publicado: (2025)
por: Wu, Jialong, et al.
Publicado: (2025)
Training Data Selection with Gradient Orthogonality for Efficient Domain Adaptation
por: Zhang, Xiyang, et al.
Publicado: (2026)
por: Zhang, Xiyang, et al.
Publicado: (2026)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
por: Ye, Hao, et al.
Publicado: (2026)
por: Ye, Hao, et al.
Publicado: (2026)
Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
por: Samineni, Soumya Rani, et al.
Publicado: (2025)
por: Samineni, Soumya Rani, et al.
Publicado: (2025)
RLPR: Extrapolating RLVR to General Domains without Verifiers
por: Yu, Tianyu, et al.
Publicado: (2025)
por: Yu, Tianyu, et al.
Publicado: (2025)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
por: Xiong, Zidi, et al.
Publicado: (2026)
por: Xiong, Zidi, et al.
Publicado: (2026)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
por: Miao, Yuchun, et al.
Publicado: (2026)
por: Miao, Yuchun, et al.
Publicado: (2026)
Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs
por: Khosravi, Hamed, et al.
Publicado: (2026)
por: Khosravi, Hamed, et al.
Publicado: (2026)
GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR
por: Zhang, Jiaying, et al.
Publicado: (2026)
por: Zhang, Jiaying, et al.
Publicado: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
por: Wu, Junkang, et al.
Publicado: (2025)
por: Wu, Junkang, et al.
Publicado: (2025)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
por: Cui, Sijia, et al.
Publicado: (2026)
por: Cui, Sijia, et al.
Publicado: (2026)
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
por: Mou, Chaoli, et al.
Publicado: (2026)
por: Mou, Chaoli, et al.
Publicado: (2026)
Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
por: Zhang, Mozhi, et al.
Publicado: (2025)
por: Zhang, Mozhi, et al.
Publicado: (2025)
The Path Not Taken: RLVR Provably Learns Off the Principals
por: Zhu, Hanqing, et al.
Publicado: (2025)
por: Zhu, Hanqing, et al.
Publicado: (2025)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
por: Ge, Albert, et al.
Publicado: (2025)
por: Ge, Albert, et al.
Publicado: (2025)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
por: Chen, Kun, et al.
Publicado: (2026)
por: Chen, Kun, et al.
Publicado: (2026)
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
por: Bauer, Justin, et al.
Publicado: (2026)
por: Bauer, Justin, et al.
Publicado: (2026)
Learning from Streaming Data when Users Choose
por: Su, Jinyan, et al.
Publicado: (2024)
por: Su, Jinyan, et al.
Publicado: (2024)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
por: Hao, Zhezheng, et al.
Publicado: (2025)
por: Hao, Zhezheng, et al.
Publicado: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
por: Duo, Jiangshan, et al.
Publicado: (2026)
por: Duo, Jiangshan, et al.
Publicado: (2026)
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
por: Qu, Yun, et al.
Publicado: (2026)
por: Qu, Yun, et al.
Publicado: (2026)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
por: Huang, Zhuoxu, et al.
Publicado: (2026)
por: Huang, Zhuoxu, et al.
Publicado: (2026)
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
por: Mitsuhashi, Ryo, et al.
Publicado: (2026)
por: Mitsuhashi, Ryo, et al.
Publicado: (2026)
The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
por: Yao, Xinhao, et al.
Publicado: (2025)
por: Yao, Xinhao, et al.
Publicado: (2025)
Deep Learning for Cross-Domain Data Fusion in Urban Computing: Taxonomy, Advances, and Outlook
por: Zou, Xingchen, et al.
Publicado: (2024)
por: Zou, Xingchen, et al.
Publicado: (2024)
Adapting Amidst Degradation: Cross Domain Li-ion Battery Health Estimation via Physics-Guided Test-Time Training
por: Feng, Yuyuan, et al.
Publicado: (2024)
por: Feng, Yuyuan, et al.
Publicado: (2024)
Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
por: He, Yuhang, et al.
Publicado: (2026)
por: He, Yuhang, et al.
Publicado: (2026)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
por: Huang, Kexin, et al.
Publicado: (2026)
por: Huang, Kexin, et al.
Publicado: (2026)
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
por: Liu, Zeyu, et al.
Publicado: (2025)
por: Liu, Zeyu, et al.
Publicado: (2025)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
por: Helff, Lukas, et al.
Publicado: (2026)
por: Helff, Lukas, et al.
Publicado: (2026)
The Multiple Ticket Hypothesis: Random Sparse Subnetworks Suffice for RLVR
por: Adewuyi, Israel, et al.
Publicado: (2026)
por: Adewuyi, Israel, et al.
Publicado: (2026)
Data-Efficient Training by Evolved Sampling
por: Cheng, Ziheng, et al.
Publicado: (2025)
por: Cheng, Ziheng, et al.
Publicado: (2025)
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
por: Shu, Yao, et al.
Publicado: (2026)
por: Shu, Yao, et al.
Publicado: (2026)
Ejemplares similares
-
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
por: Liu, Zhanyu, et al.
Publicado: (2026) -
VL Norm: Rethink Loss Aggregation in RLVR
por: He, Zhiyuan, et al.
Publicado: (2025) -
Choosing How to Remember: Adaptive Memory Structures for LLM Agents
por: Lu, Mingfei, et al.
Publicado: (2026) -
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
por: Lu, Han, et al.
Publicado: (2025) -
Spurious Rewards: Rethinking Training Signals in RLVR
por: Shao, Rulin, et al.
Publicado: (2025)