Data-Efficient RLVR via Off-Policy Influence Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Erle, Jiang, Dazhi, Wang, Yuan, Li, Xujun, Cheng, Jiale, Gu, Yuxian, Niu, Yilin, Zeng, Aohan, Tang, Jie, Huang, Minlie, Wang, Hongning |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025)
by: Su, Mingyu, et al.
Published: (2025)
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
by: Hou, Zhenyu, et al.
Published: (2024)
by: Hou, Zhenyu, et al.
Published: (2024)
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
by: Wen, Bosi, et al.
Published: (2025)
by: Wen, Bosi, et al.
Published: (2025)
Towards Efficient Exact Optimization of Language Model Alignment
by: Ji, Haozhe, et al.
Published: (2024)
by: Ji, Haozhe, et al.
Published: (2024)
MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science
by: Zhu, Erle, et al.
Published: (2025)
by: Zhu, Erle, et al.
Published: (2025)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
by: Hou, Zhenyu, et al.
Published: (2024)
by: Hou, Zhenyu, et al.
Published: (2024)
Data Selection via Optimal Control for Language Models
by: Gu, Yuxian, et al.
Published: (2024)
by: Gu, Yuxian, et al.
Published: (2024)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
by: Ke, Pei, et al.
Published: (2023)
by: Ke, Pei, et al.
Published: (2023)
MiniLLM: On-Policy Distillation of Large Language Models
by: Gu, Yuxian, et al.
Published: (2023)
by: Gu, Yuxian, et al.
Published: (2023)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
by: Cheng, Jiale, et al.
Published: (2024)
by: Cheng, Jiale, et al.
Published: (2024)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
by: Gui, Jiayi, et al.
Published: (2024)
by: Gui, Jiayi, et al.
Published: (2024)
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
by: Cheng, Jiale, et al.
Published: (2023)
by: Cheng, Jiale, et al.
Published: (2023)
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
by: Ren, Yanwei, et al.
Published: (2026)
by: Ren, Yanwei, et al.
Published: (2026)
LongSafety: Evaluating Long-Context Safety of Large Language Models
by: Lu, Yida, et al.
Published: (2025)
by: Lu, Yida, et al.
Published: (2025)
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
by: Cheng, Jiale, et al.
Published: (2024)
by: Cheng, Jiale, et al.
Published: (2024)
Evaluating Parameter Efficient Methods for RLVR
by: Yin, Qingyu, et al.
Published: (2025)
by: Yin, Qingyu, et al.
Published: (2025)
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR
by: Nayak, Anupam, et al.
Published: (2026)
by: Nayak, Anupam, et al.
Published: (2026)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
by: Wen, Bosi, et al.
Published: (2026)
by: Wen, Bosi, et al.
Published: (2026)
MiniPLM: Knowledge Distillation for Pre-Training Language Models
by: Gu, Yuxian, et al.
Published: (2024)
by: Gu, Yuxian, et al.
Published: (2024)
HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
by: Feng, Andrew Zhuoer, et al.
Published: (2026)
by: Feng, Andrew Zhuoer, et al.
Published: (2026)
Language Model Decoding as Direct Metrics Optimization
by: Ji, Haozhe, et al.
Published: (2023)
by: Ji, Haozhe, et al.
Published: (2023)
SkillEvolver: Skill Learning as a Meta-Skill
by: Zhang, Genrui, et al.
Published: (2026)
by: Zhang, Genrui, et al.
Published: (2026)
Learning to Reason under Off-Policy Guidance
by: Yan, Jianhao, et al.
Published: (2025)
by: Yan, Jianhao, et al.
Published: (2025)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators
by: Wen, Bosi, et al.
Published: (2025)
by: Wen, Bosi, et al.
Published: (2025)
Instruction Pre-Training: Language Models are Supervised Multitask Learners
by: Cheng, Daixuan, et al.
Published: (2024)
by: Cheng, Daixuan, et al.
Published: (2024)
Understanding Emergent Abilities of Language Models from the Loss Perspective
by: Du, Zhengxiao, et al.
Published: (2024)
by: Du, Zhengxiao, et al.
Published: (2024)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
by: Zhao, Jiale, et al.
Published: (2026)
by: Zhao, Jiale, et al.
Published: (2026)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
by: Yang, Junxiao, et al.
Published: (2025)
by: Yang, Junxiao, et al.
Published: (2025)
Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
by: Cui, Sijia, et al.
Published: (2026)
by: Cui, Sijia, et al.
Published: (2026)
Glyph: Scaling Context Windows via Visual-Text Compression
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
by: Zeng, Aohan, et al.
Published: (2024)
by: Zeng, Aohan, et al.
Published: (2024)
The Path Not Taken: RLVR Provably Learns Off the Principals
by: Zhu, Hanqing, et al.
Published: (2025)
by: Zhu, Hanqing, et al.
Published: (2025)
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
by: Min, Zijun, et al.
Published: (2026)
by: Min, Zijun, et al.
Published: (2026)
Towards Optimal Learning of Language Models
by: Gu, Yuxian, et al.
Published: (2024)
by: Gu, Yuxian, et al.
Published: (2024)
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
by: Zhang, Zhexin, et al.
Published: (2025)
by: Zhang, Zhexin, et al.
Published: (2025)
APAR: LLMs Can Do Auto-Parallel Auto-Regressive Decoding
by: Liu, Mingdao, et al.
Published: (2024)
by: Liu, Mingdao, et al.
Published: (2024)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
Similar Items
-
Trust-Region Adaptive Policy Optimization
by: Su, Mingyu, et al.
Published: (2025) -
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
by: Hou, Zhenyu, et al.
Published: (2024) -
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
by: Wen, Bosi, et al.
Published: (2025) -
Towards Efficient Exact Optimization of Language Model Alignment
by: Ji, Haozhe, et al.
Published: (2024) -
MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science
by: Zhu, Erle, et al.
Published: (2025)