Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
Fuente:
arXiv
Saved in:
| Main Authors: | Lochab, Anamika, Li, Bolian, Zhang, Ruqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Energy-Based Reward Models for Robust Language Model Alignment
by: Lochab, Anamika, et al.
Published: (2025)
by: Lochab, Anamika, et al.
Published: (2025)
Cascade Reward Sampling for Efficient Decoding-Time Alignment
by: Li, Bolian, et al.
Published: (2024)
by: Li, Bolian, et al.
Published: (2024)
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
by: Li, Bolian, et al.
Published: (2026)
by: Li, Bolian, et al.
Published: (2026)
VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
by: Liao, Qilin, et al.
Published: (2025)
by: Liao, Qilin, et al.
Published: (2025)
VERA: Variational Inference Framework for Jailbreaking Large Language Models
by: Lochab, Anamika, et al.
Published: (2025)
by: Lochab, Anamika, et al.
Published: (2025)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
by: Ding, Yi, et al.
Published: (2026)
by: Ding, Yi, et al.
Published: (2026)
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
by: Ding, Yi, et al.
Published: (2024)
by: Ding, Yi, et al.
Published: (2024)
Entropy-MCMC: Sampling from Flat Basins with Ease
by: Li, Bolian, et al.
Published: (2023)
by: Li, Bolian, et al.
Published: (2023)
Making Reliable and Flexible Decisions in Long-tailed Classification
by: Li, Bolian, et al.
Published: (2025)
by: Li, Bolian, et al.
Published: (2025)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
by: Cui, Sijia, et al.
Published: (2026)
by: Cui, Sijia, et al.
Published: (2026)
Sherlock: Self-Correcting Reasoning in Vision-Language Models
by: Ding, Yi, et al.
Published: (2025)
by: Ding, Yi, et al.
Published: (2025)
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
by: Heakl, Ahmed, et al.
Published: (2026)
by: Heakl, Ahmed, et al.
Published: (2026)
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
Why Any-Order Autoregressive Models Need Two-Stream Attention: A Structural-Semantic Tradeoff
by: Pynadath, Patrick, et al.
Published: (2026)
by: Pynadath, Patrick, et al.
Published: (2026)
Controlled LLM Decoding via Discrete Auto-regressive Biasing
by: Pynadath, Patrick, et al.
Published: (2025)
by: Pynadath, Patrick, et al.
Published: (2025)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
The Unlearnability Phenomenon in RLVR for Language Models
by: Chen, Yulin, et al.
Published: (2026)
by: Chen, Yulin, et al.
Published: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
by: He, Bingxiang, et al.
Published: (2026)
by: He, Bingxiang, et al.
Published: (2026)
Generative Frontiers: Why Evaluation Matters for Diffusion Language Models
by: Pynadath, Patrick, et al.
Published: (2026)
by: Pynadath, Patrick, et al.
Published: (2026)
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
by: Li, Bolian, et al.
Published: (2025)
by: Li, Bolian, et al.
Published: (2025)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
by: Chen, Zhipeng, et al.
Published: (2026)
by: Chen, Zhipeng, et al.
Published: (2026)
Bayesian Computation in Deep Learning
by: Chen, Wenlong, et al.
Published: (2025)
by: Chen, Wenlong, et al.
Published: (2025)
Efficient RLVR Training via Weighted Mutual Information Data Selection
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
by: Khandoga, Mykola, et al.
Published: (2026)
by: Khandoga, Mykola, et al.
Published: (2026)
Linear Dynamics in the RLVR Training of Large Language Models
by: Wang, Tianle, et al.
Published: (2026)
by: Wang, Tianle, et al.
Published: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
by: Duo, Jiangshan, et al.
Published: (2026)
by: Duo, Jiangshan, et al.
Published: (2026)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
by: Zhai, Zepeng, et al.
Published: (2026)
by: Zhai, Zepeng, et al.
Published: (2026)
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
by: Zhao, Jiale, et al.
Published: (2026)
by: Zhao, Jiale, et al.
Published: (2026)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
by: Li, Kaiyuan, et al.
Published: (2026)
by: Li, Kaiyuan, et al.
Published: (2026)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
by: Zhang, Zhaowei, et al.
Published: (2026)
by: Zhang, Zhaowei, et al.
Published: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
by: Gu, Hengrui, et al.
Published: (2026)
by: Gu, Hengrui, et al.
Published: (2026)
RLPR: Extrapolating RLVR to General Domains without Verifiers
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
Generalization of RLVR Using Causal Reasoning as a Testbed
by: Lu, Brian, et al.
Published: (2025)
by: Lu, Brian, et al.
Published: (2025)
Breaking MLPerf Training: A Case Study on Optimizing BERT
by: Kim, Yongdeok, et al.
Published: (2024)
by: Kim, Yongdeok, et al.
Published: (2024)
Similar Items
-
Energy-Based Reward Models for Robust Language Model Alignment
by: Lochab, Anamika, et al.
Published: (2025) -
Cascade Reward Sampling for Efficient Decoding-Time Alignment
by: Li, Bolian, et al.
Published: (2024) -
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
by: Li, Bolian, et al.
Published: (2026) -
VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
by: Liao, Qilin, et al.
Published: (2025) -
VERA: Variational Inference Framework for Jailbreaking Large Language Models
by: Lochab, Anamika, et al.
Published: (2025)