Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jiakang, Liu, Runze, Zhang, Fuzheng, Li, Xiu, Zhou, Guorui, Pan, Ling |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
by: Wang, Jiakang, et al.
Published: (2025)
by: Wang, Jiakang, et al.
Published: (2025)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
by: Yuan, Zhonghang, et al.
Published: (2026)
by: Yuan, Zhonghang, et al.
Published: (2026)
Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR
by: Zhuang, Haomin, et al.
Published: (2025)
by: Zhuang, Haomin, et al.
Published: (2025)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
by: Duo, Jiangshan, et al.
Published: (2026)
by: Duo, Jiangshan, et al.
Published: (2026)
Character-R1: Enhancing Role-Aware Reasoning in Role-Playing Agents via RLVR
by: Tang, Yihong, et al.
Published: (2026)
by: Tang, Yihong, et al.
Published: (2026)
Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
by: Ji, Xingguang, et al.
Published: (2025)
by: Ji, Xingguang, et al.
Published: (2025)
Detecting RLVR Training Data via Structural Convergence of Reasoning
by: Zhang, Hongbo, et al.
Published: (2026)
by: Zhang, Hongbo, et al.
Published: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
by: Zhao, Jian, et al.
Published: (2025)
by: Zhao, Jian, et al.
Published: (2025)
Data Metabolism: An Efficient Data Design Schema For Vision Language Model
by: Zhang, Jingyuan, et al.
Published: (2025)
by: Zhang, Jingyuan, et al.
Published: (2025)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
by: Wang, Zhilin, et al.
Published: (2026)
by: Wang, Zhilin, et al.
Published: (2026)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
by: Meng, Haoming, et al.
Published: (2026)
by: Meng, Haoming, et al.
Published: (2026)
Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint
by: Chen, Zhipeng, et al.
Published: (2024)
by: Chen, Zhipeng, et al.
Published: (2024)
Dual Reasoning: A GNN-LLM Collaborative Framework for Knowledge Graph Question Answering
by: Liu, Guangyi, et al.
Published: (2024)
by: Liu, Guangyi, et al.
Published: (2024)
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models
by: Yi, Hao, et al.
Published: (2024)
by: Yi, Hao, et al.
Published: (2024)
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
by: Zhu, Yihua, et al.
Published: (2026)
by: Zhu, Yihua, et al.
Published: (2026)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Tokenization Constraints in LLMs: A Study of Symbolic and Arithmetic Reasoning Limits
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Generalization of RLVR Using Causal Reasoning as a Testbed
by: Lu, Brian, et al.
Published: (2025)
by: Lu, Brian, et al.
Published: (2025)
Efficient RLVR Training via Weighted Mutual Information Data Selection
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
A State-of-the-Art SQL Reasoning Model using RLVR
by: Ali, Alnur, et al.
Published: (2025)
by: Ali, Alnur, et al.
Published: (2025)
Knowledge Reasoning Language Model: Unifying Knowledge and Language for Inductive Knowledge Graph Reasoning
by: Zhuo, Xingrui, et al.
Published: (2025)
by: Zhuo, Xingrui, et al.
Published: (2025)
LLM-Driven Reasoning for Constraint-Aware Feature Selection in Industrial Systems
by: Zhou, Yuhang, et al.
Published: (2026)
by: Zhou, Yuhang, et al.
Published: (2026)
KnowledgeNavigator: Leveraging Large Language Models for Enhanced Reasoning over Knowledge Graph
by: Guo, Tiezheng, et al.
Published: (2023)
by: Guo, Tiezheng, et al.
Published: (2023)
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
by: Zhang, Sheng, et al.
Published: (2025)
by: Zhang, Sheng, et al.
Published: (2025)
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Document Reconstruction Unlocks Scalable Long-Context RLVR
by: Xiao, Yao, et al.
Published: (2026)
by: Xiao, Yao, et al.
Published: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
by: He, Bingxiang, et al.
Published: (2026)
by: He, Bingxiang, et al.
Published: (2026)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
by: Kim, Minsang, et al.
Published: (2026)
by: Kim, Minsang, et al.
Published: (2026)
TwT: Thinking without Tokens by Habitual Reasoning Distillation with Multi-Teachers' Guidance
by: Xu, Jingxian, et al.
Published: (2025)
by: Xu, Jingxian, et al.
Published: (2025)
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
by: Kim, Soeun, et al.
Published: (2026)
by: Kim, Soeun, et al.
Published: (2026)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
by: Lochab, Anamika, et al.
Published: (2026)
by: Lochab, Anamika, et al.
Published: (2026)
mAceReason-Math: A Dataset of High-Quality Multilingual Math Problems Ready For RLVR
by: Dobler, Konstantin, et al.
Published: (2026)
by: Dobler, Konstantin, et al.
Published: (2026)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
Similar Items
-
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
by: Wang, Jiakang, et al.
Published: (2025) -
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025) -
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
by: Zhang, Hongzhi, et al.
Published: (2025) -
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
by: Yuan, Zhonghang, et al.
Published: (2026) -
Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR
by: Zhuang, Haomin, et al.
Published: (2025)