Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Pei-Xi, Lin, Che-Yu, Yang, Cheng-Lin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning to Hint for Reinforcement Learning
von: Xia, Yu, et al.
Veröffentlicht: (2026)
von: Xia, Yu, et al.
Veröffentlicht: (2026)
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2025)
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2025)
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
von: Li, Pingzhi, et al.
Veröffentlicht: (2023)
von: Li, Pingzhi, et al.
Veröffentlicht: (2023)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
Self-Hinting Language Models Enhance Reinforcement Learning
von: Liao, Baohao, et al.
Veröffentlicht: (2026)
von: Liao, Baohao, et al.
Veröffentlicht: (2026)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
von: Lee, Heekyung, et al.
Veröffentlicht: (2025)
von: Lee, Heekyung, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2026)
von: Yu, Yongcan, et al.
Veröffentlicht: (2026)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
von: Tang, Siao, et al.
Veröffentlicht: (2025)
von: Tang, Siao, et al.
Veröffentlicht: (2025)
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
von: Wang, Shengyuan, et al.
Veröffentlicht: (2025)
von: Wang, Shengyuan, et al.
Veröffentlicht: (2025)
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
von: Yang, An, et al.
Veröffentlicht: (2024)
von: Yang, An, et al.
Veröffentlicht: (2024)
Progressive-Hint Prompting Improves Reasoning in Large Language Models
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2023)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2023)
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
von: Li, Junsong, et al.
Veröffentlicht: (2025)
von: Li, Junsong, et al.
Veröffentlicht: (2025)
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
von: Tran, Thien Q., et al.
Veröffentlicht: (2025)
von: Tran, Thien Q., et al.
Veröffentlicht: (2025)
Self-Improvement in Language Models: The Sharpening Mechanism
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
MegaMath: Pushing the Limits of Open Math Corpora
von: Zhou, Fan, et al.
Veröffentlicht: (2025)
von: Zhou, Fan, et al.
Veröffentlicht: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
RLPR: Extrapolating RLVR to General Domains without Verifiers
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
Finetune Once: Decoupling General & Domain Learning with Dynamic Boosted Annealing
von: Tang, Yang, et al.
Veröffentlicht: (2025)
von: Tang, Yang, et al.
Veröffentlicht: (2025)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
von: Chen, Kun, et al.
Veröffentlicht: (2026)
von: Chen, Kun, et al.
Veröffentlicht: (2026)
Fill in the Blank: Exploring and Enhancing LLM Capabilities for Backward Reasoning in Math Word Problems
von: Deb, Aniruddha, et al.
Veröffentlicht: (2023)
von: Deb, Aniruddha, et al.
Veröffentlicht: (2023)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
Generalization of RLVR Using Causal Reasoning as a Testbed
von: Lu, Brian, et al.
Veröffentlicht: (2025)
von: Lu, Brian, et al.
Veröffentlicht: (2025)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline
von: Seegmiller, Parker, et al.
Veröffentlicht: (2025)
von: Seegmiller, Parker, et al.
Veröffentlicht: (2025)
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
von: Luo, Haipeng, et al.
Veröffentlicht: (2023)
von: Luo, Haipeng, et al.
Veröffentlicht: (2023)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
Forward-Backward Reasoning in Large Language Models for Mathematical Verification
von: Jiang, Weisen, et al.
Veröffentlicht: (2023)
von: Jiang, Weisen, et al.
Veröffentlicht: (2023)
Lexical Hints of Accuracy in LLM Reasoning Chains
von: Vanhoyweghen, Arne, et al.
Veröffentlicht: (2025)
von: Vanhoyweghen, Arne, et al.
Veröffentlicht: (2025)
The Invisible Leash: Why RLVR May or May Not Escape Its Origin
von: Wu, Fang, et al.
Veröffentlicht: (2025)
von: Wu, Fang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning to Hint for Reinforcement Learning
von: Xia, Yu, et al.
Veröffentlicht: (2026) -
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2025) -
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
von: Li, Pingzhi, et al.
Veröffentlicht: (2023) -
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
von: Xie, Yanyue, et al.
Veröffentlicht: (2024) -
Self-Hinting Language Models Enhance Reinforcement Learning
von: Liao, Baohao, et al.
Veröffentlicht: (2026)