Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Hengrui, Han, Xiaotian, Bian, Yujing, Wang, Feiyi, Zhou, Kaixiong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Domains Interact: Asymmetric and Order-Sensitive Cross-Domain Effects in Reinforcement Learning for Reasoning
by: Yang, Wang, et al.
Published: (2026)
by: Yang, Wang, et al.
Published: (2026)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
by: Chen, Kun, et al.
Published: (2026)
by: Chen, Kun, et al.
Published: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Pioneering Reliable Assessment in Text-to-Image Knowledge Editing: Leveraging a Fine-Grained Dataset and an Innovative Criterion
by: Gu, Hengrui, et al.
Published: (2024)
by: Gu, Hengrui, et al.
Published: (2024)
On the Entropy Calibration of Language Models
by: Cao, Steven, et al.
Published: (2025)
by: Cao, Steven, et al.
Published: (2025)
RLPR: Extrapolating RLVR to General Domains without Verifiers
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages
by: Kunde, Vishnu Teja, et al.
Published: (2026)
by: Kunde, Vishnu Teja, et al.
Published: (2026)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
by: Duo, Jiangshan, et al.
Published: (2026)
by: Duo, Jiangshan, et al.
Published: (2026)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
by: Meng, Haoming, et al.
Published: (2026)
by: Meng, Haoming, et al.
Published: (2026)
Cross-Lingual Multi-Hop Knowledge Editing
by: Khandelwal, Aditi, et al.
Published: (2024)
by: Khandelwal, Aditi, et al.
Published: (2024)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
by: Cui, Sijia, et al.
Published: (2026)
by: Cui, Sijia, et al.
Published: (2026)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Generalization of RLVR Using Causal Reasoning as a Testbed
by: Lu, Brian, et al.
Published: (2025)
by: Lu, Brian, et al.
Published: (2025)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
by: Khalifa, Muhammad, et al.
Published: (2026)
by: Khalifa, Muhammad, et al.
Published: (2026)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
by: Chen, Zhipeng, et al.
Published: (2026)
by: Chen, Zhipeng, et al.
Published: (2026)
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
by: Li, Sijia, et al.
Published: (2026)
by: Li, Sijia, et al.
Published: (2026)
Retrieval-enhanced Knowledge Editing in Language Models for Multi-Hop Question Answering
by: Shi, Yucheng, et al.
Published: (2024)
by: Shi, Yucheng, et al.
Published: (2024)
The Invisible Leash: Why RLVR May or May Not Escape Its Origin
by: Wu, Fang, et al.
Published: (2025)
by: Wu, Fang, et al.
Published: (2025)
Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering
by: Zhou, Han, et al.
Published: (2023)
by: Zhou, Han, et al.
Published: (2023)
MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning
by: Wang, Hongjun, et al.
Published: (2026)
by: Wang, Hongjun, et al.
Published: (2026)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
by: Wang, Cheng, et al.
Published: (2026)
by: Wang, Cheng, et al.
Published: (2026)
No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
by: Le, Thanh-Long V., et al.
Published: (2025)
by: Le, Thanh-Long V., et al.
Published: (2025)
Text Rationalization for Robust Causal Effect Estimation
by: Zhang, Lijinghua, et al.
Published: (2025)
by: Zhang, Lijinghua, et al.
Published: (2025)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
by: Kim, Soeun, et al.
Published: (2026)
by: Kim, Soeun, et al.
Published: (2026)
A State-of-the-Art SQL Reasoning Model using RLVR
by: Ali, Alnur, et al.
Published: (2025)
by: Ali, Alnur, et al.
Published: (2025)
Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thought
by: Chen, Zhikang, et al.
Published: (2025)
by: Chen, Zhikang, et al.
Published: (2025)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
by: Wang, Shouren, et al.
Published: (2025)
by: Wang, Shouren, et al.
Published: (2025)
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
by: Xie, Pei-Xi, et al.
Published: (2026)
by: Xie, Pei-Xi, et al.
Published: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
by: Zhang, Wenbo, et al.
Published: (2026)
by: Zhang, Wenbo, et al.
Published: (2026)
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
by: Cai, Hengrui, et al.
Published: (2023)
by: Cai, Hengrui, et al.
Published: (2023)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
by: Hao, Zhezheng, et al.
Published: (2025)
by: Hao, Zhezheng, et al.
Published: (2025)
Hard Prompts Made Interpretable: Sparse Entropy Regularization for Prompt Tuning with RL
by: Choi, Yunseon, et al.
Published: (2024)
by: Choi, Yunseon, et al.
Published: (2024)
Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation
by: Akarlar, G. Aytug
Published: (2026)
by: Akarlar, G. Aytug
Published: (2026)
Entropy Adaptive Decoding: Dynamic Model Switching for Efficient Inference
by: Simonds, Toby
Published: (2025)
by: Simonds, Toby
Published: (2025)
iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use
by: Zeng, Yirong, et al.
Published: (2025)
by: Zeng, Yirong, et al.
Published: (2025)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
by: Maheswaran, Monishwaran, et al.
Published: (2025)
by: Maheswaran, Monishwaran, et al.
Published: (2025)
Similar Items
-
When Domains Interact: Asymmetric and Order-Sensitive Cross-Domain Effects in Reinforcement Learning for Reasoning
by: Yang, Wang, et al.
Published: (2026) -
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
by: Chen, Kun, et al.
Published: (2026) -
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025) -
Pioneering Reliable Assessment in Text-to-Image Knowledge Editing: Leveraging a Fine-Grained Dataset and an Innovative Criterion
by: Gu, Hengrui, et al.
Published: (2024) -
On the Entropy Calibration of Language Models
by: Cao, Steven, et al.
Published: (2025)