Efficient RLVR Training via Weighted Mutual Information Data Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Xinyu, Zhu, Boyu, Zhang, Haotian, Wang, Huiming, Guo, Zhijiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
Balancing the Reasoning Load: Difficulty-Differentiated Policy Optimization with Length Redistribution for Efficient and Robust Reinforcement Learning
von: Xia, Yinan, et al.
Veröffentlicht: (2026)
von: Xia, Yinan, et al.
Veröffentlicht: (2026)
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
von: Wei, Zhepei, et al.
Veröffentlicht: (2026)
von: Wei, Zhepei, et al.
Veröffentlicht: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
Linear Dynamics in the RLVR Training of Large Language Models
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning
von: Deng, Mengyi, et al.
Veröffentlicht: (2025)
von: Deng, Mengyi, et al.
Veröffentlicht: (2025)
Self-Distilled RLVR
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
Scaling Multi-Hop Training Data via Graph-Constrained Path Selection
von: Chen, Pengyu, et al.
Veröffentlicht: (2026)
von: Chen, Pengyu, et al.
Veröffentlicht: (2026)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
Efficient Process Reward Modeling via Contrastive Mutual Information
von: Lee, Nakyung, et al.
Veröffentlicht: (2026)
von: Lee, Nakyung, et al.
Veröffentlicht: (2026)
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
von: Yao, Jiashu, et al.
Veröffentlicht: (2026)
von: Yao, Jiashu, et al.
Veröffentlicht: (2026)
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
The Unlearnability Phenomenon in RLVR for Language Models
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
von: Xu, Minrui, et al.
Veröffentlicht: (2026)
von: Xu, Minrui, et al.
Veröffentlicht: (2026)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
von: Lochab, Anamika, et al.
Veröffentlicht: (2026)
von: Lochab, Anamika, et al.
Veröffentlicht: (2026)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
DavIR: Data Selection via Implicit Reward for Large Language Models
von: Zhou, Haotian, et al.
Veröffentlicht: (2023)
von: Zhou, Haotian, et al.
Veröffentlicht: (2023)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
von: Li, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Li, Kaiyuan, et al.
Veröffentlicht: (2026)
Efficient Data Selection at Scale via Influence Distillation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
Studying the Korean Word-Chain Game with RLVR: Mitigating Reward Conflicts via Curriculum Learning
von: Rho, Donghwan
Veröffentlicht: (2025)
von: Rho, Donghwan
Veröffentlicht: (2025)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
von: Zhai, Zepeng, et al.
Veröffentlicht: (2026)
von: Zhai, Zepeng, et al.
Veröffentlicht: (2026)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
RLPR: Extrapolating RLVR to General Domains without Verifiers
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
von: Xu, Haotian, et al.
Veröffentlicht: (2025)
von: Xu, Haotian, et al.
Veröffentlicht: (2025)
Intrinsic Mutual Information as a Modulator for Preference Optimization
von: Liao, Peng, et al.
Veröffentlicht: (2026)
von: Liao, Peng, et al.
Veröffentlicht: (2026)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
Skip-Connected Policy Optimization for Implicit Advantage
von: Teng, Fengwei, et al.
Veröffentlicht: (2026)
von: Teng, Fengwei, et al.
Veröffentlicht: (2026)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
von: Fan, Ziqing, et al.
Veröffentlicht: (2025)
von: Fan, Ziqing, et al.
Veröffentlicht: (2025)
Group-Level Data Selection for Efficient Pretraining
von: Yu, Zichun, et al.
Veröffentlicht: (2025)
von: Yu, Zichun, et al.
Veröffentlicht: (2025)
Influence Functions for Efficient Data Selection in Reasoning
von: Humane, Prateek, et al.
Veröffentlicht: (2025)
von: Humane, Prateek, et al.
Veröffentlicht: (2025)
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
von: Zhu, Dawei, et al.
Veröffentlicht: (2023)
von: Zhu, Dawei, et al.
Veröffentlicht: (2023)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
von: Li, Jiacheng, et al.
Veröffentlicht: (2025)
von: Li, Jiacheng, et al.
Veröffentlicht: (2025)
From the Inside Out: Progressive Distribution Refinement for Confidence Calibration
von: Yang, Xizhong, et al.
Veröffentlicht: (2026)
von: Yang, Xizhong, et al.
Veröffentlicht: (2026)
AutoPSV: Automated Process-Supervised Verifier
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
von: Zhu, Xiao, et al.
Veröffentlicht: (2026) -
Balancing the Reasoning Load: Difficulty-Differentiated Policy Optimization with Length Redistribution for Efficient and Robust Reinforcement Learning
von: Xia, Yinan, et al.
Veröffentlicht: (2026) -
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
von: Wei, Zhepei, et al.
Veröffentlicht: (2026) -
How Far Can Unsupervised RLVR Scale LLM Training?
von: He, Bingxiang, et al.
Veröffentlicht: (2026) -
Linear Dynamics in the RLVR Training of Large Language Models
von: Wang, Tianle, et al.
Veröffentlicht: (2026)