Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Yanwei, Zhang, Haotian, Xiao, Likang, Zhang, Xikai, Huang, Jiaxing, Qiu, Jiayan, Yu, Baosheng, Chen, Quan, Liu, Liu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
by: Zhang, Haotian, et al.
Published: (2025)
by: Zhang, Haotian, et al.
Published: (2025)
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
by: Cui, Yitong, et al.
Published: (2025)
by: Cui, Yitong, et al.
Published: (2025)
SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation
by: Ren, Yanwei, et al.
Published: (2025)
by: Ren, Yanwei, et al.
Published: (2025)
LARGO: Low-Rank Regulated Gradient Projection for Robust Parameter Efficient Fine-Tuning
by: Zhang, Haotian, et al.
Published: (2025)
by: Zhang, Haotian, et al.
Published: (2025)
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
by: Ren, Yanwei, et al.
Published: (2025)
by: Ren, Yanwei, et al.
Published: (2025)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
Data-Efficient RLVR via Off-Policy Influence Guidance
by: Zhu, Erle, et al.
Published: (2025)
by: Zhu, Erle, et al.
Published: (2025)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
by: Wu, Xiangyang, et al.
Published: (2025)
by: Wu, Xiangyang, et al.
Published: (2025)
IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
by: Zhang, Xikai, et al.
Published: (2025)
by: Zhang, Xikai, et al.
Published: (2025)
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning
by: Zhang, Xikai, et al.
Published: (2026)
by: Zhang, Xikai, et al.
Published: (2026)
Re-Initialization Token Learning for Tool-Augmented Large Language Models
by: Li, Chenghao, et al.
Published: (2025)
by: Li, Chenghao, et al.
Published: (2025)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
by: Huang, Zhuoxu, et al.
Published: (2026)
by: Huang, Zhuoxu, et al.
Published: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
by: Xu, Haohang, et al.
Published: (2026)
by: Xu, Haohang, et al.
Published: (2026)
The Path Not Taken: RLVR Provably Learns Off the Principals
by: Zhu, Hanqing, et al.
Published: (2025)
by: Zhu, Hanqing, et al.
Published: (2025)
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR
by: Nayak, Anupam, et al.
Published: (2026)
by: Nayak, Anupam, et al.
Published: (2026)
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
by: Liu, Zhanyu, et al.
Published: (2026)
by: Liu, Zhanyu, et al.
Published: (2026)
Efficient RLVR Training via Weighted Mutual Information Data Selection
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Learning to Reason under Off-Policy Guidance
by: Yan, Jianhao, et al.
Published: (2025)
by: Yan, Jianhao, et al.
Published: (2025)
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
by: Min, Zijun, et al.
Published: (2026)
by: Min, Zijun, et al.
Published: (2026)
SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and More
by: Huang, Muye, et al.
Published: (2026)
by: Huang, Muye, et al.
Published: (2026)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
by: Li, Xiaofan, et al.
Published: (2026)
by: Li, Xiaofan, et al.
Published: (2026)
Cryobiopsy as a Salvage Technique Following Negative Flexible Forceps Biopsy of the Pleura Under Rapid On‐Site Evaluation Guidance: A Prospective Study
by: Jianlong Tan, et al.
Published: (2025)
by: Jianlong Tan, et al.
Published: (2025)
Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
by: Liu, Fanfan, et al.
Published: (2026)
by: Liu, Fanfan, et al.
Published: (2026)
On-Off Systems with Strategic Customers
by: Sun, Yanwei, et al.
Published: (2025)
by: Sun, Yanwei, et al.
Published: (2025)
Fine-Grained Guidance for Retrievers: Leveraging LLMs' Feedback in Retrieval-Augmented Generation
by: Liu, Yuhang, et al.
Published: (2024)
by: Liu, Yuhang, et al.
Published: (2024)
Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
by: Qiu, Zhongxi, et al.
Published: (2025)
by: Qiu, Zhongxi, et al.
Published: (2025)
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
by: Fan, Chongyu, et al.
Published: (2026)
by: Fan, Chongyu, et al.
Published: (2026)
Learning Diverse Policies with Soft Self-Generated Guidance
by: Wang, Guojian, et al.
Published: (2024)
by: Wang, Guojian, et al.
Published: (2024)
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
by: Da, Jeff, et al.
Published: (2025)
by: Da, Jeff, et al.
Published: (2025)
GrainGrasp: Dexterous Grasp Generation with Fine-grained Contact Guidance
by: Zhao, Fuqiang, et al.
Published: (2024)
by: Zhao, Fuqiang, et al.
Published: (2024)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
by: Lochab, Anamika, et al.
Published: (2026)
by: Lochab, Anamika, et al.
Published: (2026)
Turning Sand to Gold: Recycling Data to Bridge On-Policy and Off-Policy Learning via Causal Bound
by: Fiskus, Tal, et al.
Published: (2025)
by: Fiskus, Tal, et al.
Published: (2025)
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
by: Xie, Can, et al.
Published: (2025)
by: Xie, Can, et al.
Published: (2025)
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Selective Off-Policy Reference Tuning with Plan Guidance
by: Le, Duc Anh, et al.
Published: (2026)
by: Le, Duc Anh, et al.
Published: (2026)
FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with Citations
by: Peng, Yixing, et al.
Published: (2025)
by: Peng, Yixing, et al.
Published: (2025)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
by: Cui, Sijia, et al.
Published: (2026)
by: Cui, Sijia, et al.
Published: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
by: He, Zhiyuan, et al.
Published: (2025)
by: He, Zhiyuan, et al.
Published: (2025)
Similar Items
-
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
by: Zhang, Haotian, et al.
Published: (2025) -
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
by: Cui, Yitong, et al.
Published: (2025) -
SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation
by: Ren, Yanwei, et al.
Published: (2025) -
LARGO: Low-Rank Regulated Gradient Projection for Robust Parameter Efficient Fine-Tuning
by: Zhang, Haotian, et al.
Published: (2025) -
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
by: Ren, Yanwei, et al.
Published: (2025)