On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Kexin, Meng, Haoming, Wu, Junkang, Lu, Jinda, Ma, Chiyu, Chen, Ziqian, Wang, Xue, Ding, Bolin, Wu, Jiancan, Wang, Xiang, He, Xiangnan, Wang, Guoyin, Zhou, Jingren |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
by: Lu, Jinda, et al.
Published: (2026)
by: Lu, Jinda, et al.
Published: (2026)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
by: Meng, Haoming, et al.
Published: (2026)
by: Meng, Haoming, et al.
Published: (2026)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
by: Huang, Kexin, et al.
Published: (2025)
by: Huang, Kexin, et al.
Published: (2025)
FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
by: Ma, Chiyu, et al.
Published: (2026)
by: Ma, Chiyu, et al.
Published: (2026)
RePO: Understanding Preference Learning Through ReLU-Based Optimization
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
by: Lu, Jinda, et al.
Published: (2026)
by: Lu, Jinda, et al.
Published: (2026)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
by: Ji, Wence, et al.
Published: (2025)
by: Ji, Wence, et al.
Published: (2025)
Unified Parameter-Efficient Unlearning for LLMs
by: Ding, Chenlu, et al.
Published: (2024)
by: Ding, Chenlu, et al.
Published: (2024)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
One-Way Policy Optimization for Self-Evolving LLMs
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
R^2-Mem: Reflective Experience for Memory Search
by: Wang, Xinyuan, et al.
Published: (2026)
by: Wang, Xinyuan, et al.
Published: (2026)
Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal Intervention
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
by: Ding, Chenlu, et al.
Published: (2025)
by: Ding, Chenlu, et al.
Published: (2025)
Reinforced Prompt Personalization for Recommendation with Large Language Models
by: Mao, Wenyu, et al.
Published: (2024)
by: Mao, Wenyu, et al.
Published: (2024)
On Negative-aware Preference Optimization for Recommendation
by: Ding, Chenlu, et al.
Published: (2025)
by: Ding, Chenlu, et al.
Published: (2025)
RosePO: Aligning LLM-based Recommenders with Human Values
by: Liao, Jiayi, et al.
Published: (2024)
by: Liao, Jiayi, et al.
Published: (2024)
Customizing Language Models with Instance-wise LoRA for Sequential Recommendation
by: Kong, Xiaoyu, et al.
Published: (2024)
by: Kong, Xiaoyu, et al.
Published: (2024)
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
by: Li, Jinghan, et al.
Published: (2026)
by: Li, Jinghan, et al.
Published: (2026)
LLaRA: Large Language-Recommendation Assistant
by: Liao, Jiayi, et al.
Published: (2023)
by: Liao, Jiayi, et al.
Published: (2023)
Addressing Missing Data Issue for Diffusion-based Recommendation
by: Mao, Wenyu, et al.
Published: (2025)
by: Mao, Wenyu, et al.
Published: (2025)
LaMP-Val: Large Language Models Empower Personalized Valuation in Auction
by: Sun, Jie, et al.
Published: (2024)
by: Sun, Jie, et al.
Published: (2024)
MiniOneRec: An Open-Source Framework for Scaling Generative Recommendation
by: Kong, Xiaoyu, et al.
Published: (2025)
by: Kong, Xiaoyu, et al.
Published: (2025)
Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model
by: Wang, Xue, et al.
Published: (2025)
by: Wang, Xue, et al.
Published: (2025)
Robust Preference Optimization via Dynamic Target Margins
by: Sun, Jie, et al.
Published: (2025)
by: Sun, Jie, et al.
Published: (2025)
Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning
by: Ding, Chenlu, et al.
Published: (2026)
by: Ding, Chenlu, et al.
Published: (2026)
Lower-Left Partial AUC: An Effective and Efficient Optimization Metric for Recommendation
by: Shi, Wentao, et al.
Published: (2024)
by: Shi, Wentao, et al.
Published: (2024)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
Leave No Patient Behind: Enhancing Medication Recommendation for Rare Disease Patients
by: Zhao, Zihao, et al.
Published: (2024)
by: Zhao, Zihao, et al.
Published: (2024)
Enhancing Temporal Sensitivity of Large Language Model for Recommendation with Counterfactual Tuning
by: Liu, Yutian, et al.
Published: (2025)
by: Liu, Yutian, et al.
Published: (2025)
Let Me Do It For You: Towards LLM Empowered Recommendation via Tool Learning
by: Zhao, Yuyue, et al.
Published: (2024)
by: Zhao, Yuyue, et al.
Published: (2024)
Boosting Few-Shot Learning via Attentive Feature Regularization
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
by: Chen, Yanxi, et al.
Published: (2023)
by: Chen, Yanxi, et al.
Published: (2023)
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
by: Xue, Qiyao, et al.
Published: (2025)
by: Xue, Qiyao, et al.
Published: (2025)
Similar Items
-
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025) -
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
by: Lu, Jinda, et al.
Published: (2026) -
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
by: Meng, Haoming, et al.
Published: (2026) -
Larger or Smaller Reward Margins to Select Preferences for Alignment?
by: Huang, Kexin, et al.
Published: (2025) -
FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization
by: Ma, Chiyu, et al.
Published: (2026)