Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Si, Shen, Peijun, Zhao, Wenhua, Zhu, Danhao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RevOrder: A Novel Method for Enhanced Arithmetic in Language Models
by: Shen, Si, et al.
Published: (2024)
by: Shen, Si, et al.
Published: (2024)
LLM-Metrics: Measuring Research Impact Through Large Language Model Memory
by: Shen, Si, et al.
Published: (2026)
by: Shen, Si, et al.
Published: (2026)
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
by: Sun, Ke, et al.
Published: (2026)
by: Sun, Ke, et al.
Published: (2026)
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
by: Li, Sijia, et al.
Published: (2026)
by: Li, Sijia, et al.
Published: (2026)
TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning
by: Pan, Muyu, et al.
Published: (2026)
by: Pan, Muyu, et al.
Published: (2026)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
by: Wang, Cheng, et al.
Published: (2026)
by: Wang, Cheng, et al.
Published: (2026)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
by: Agrawal, Aakriti, et al.
Published: (2025)
by: Agrawal, Aakriti, et al.
Published: (2025)
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning
by: Gong, Shijin, et al.
Published: (2026)
by: Gong, Shijin, et al.
Published: (2026)
Learn to Think: Bootstrapping LLM Reasoning Capability Through Graph Representation Learning
by: Gao, Hang, et al.
Published: (2025)
by: Gao, Hang, et al.
Published: (2025)
Think Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient Reasoning
by: Peng, Keqin, et al.
Published: (2026)
by: Peng, Keqin, et al.
Published: (2026)
Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
by: Zhang, Qingyang, et al.
Published: (2025)
by: Zhang, Qingyang, et al.
Published: (2025)
Thinking with Knowledge Graphs: Enhancing LLM Reasoning Through Structured Data
by: Wu, Xue, et al.
Published: (2024)
by: Wu, Xue, et al.
Published: (2024)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024)
by: Zhang, Chuheng, et al.
Published: (2024)
Efficient Reasoning with Hidden Thinking
by: Shen, Xuan, et al.
Published: (2025)
by: Shen, Xuan, et al.
Published: (2025)
Topology-Aware Dynamic Reweighting for Distribution Shifts on Graph
by: Zheng, Weihuang, et al.
Published: (2024)
by: Zheng, Weihuang, et al.
Published: (2024)
Mitigating the Noise Shift for Denoising Generative Models via Noise Awareness Guidance
by: Zhong, Jincheng, et al.
Published: (2025)
by: Zhong, Jincheng, et al.
Published: (2025)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
by: Brantley, Kianté, et al.
Published: (2025)
by: Brantley, Kianté, et al.
Published: (2025)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
Flow-of-Options: Diversified and Improved LLM Reasoning by Thinking Through Options
by: Nair, Lakshmi, et al.
Published: (2025)
by: Nair, Lakshmi, et al.
Published: (2025)
Individual Fairness Through Reweighting and Tuning
by: Mahamadou, Abdoul Jalil Djiberou, et al.
Published: (2024)
by: Mahamadou, Abdoul Jalil Djiberou, et al.
Published: (2024)
Stable Adaptive Thinking via Advantage Shaping and Length-Aware Gradient Regulation
by: Xu, Zihang, et al.
Published: (2026)
by: Xu, Zihang, et al.
Published: (2026)
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models
by: Sun, Chung-En, et al.
Published: (2025)
by: Sun, Chung-En, et al.
Published: (2025)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
by: Maheswaran, Monishwaran, et al.
Published: (2025)
by: Maheswaran, Monishwaran, et al.
Published: (2025)
Mitigating the Language Mismatch and Repetition Issues in LLM-based Machine Translation via Model Editing
by: Wang, Weichuan, et al.
Published: (2024)
by: Wang, Weichuan, et al.
Published: (2024)
Mitigating Individual Skin Tone Bias in Skin Lesion Classification through Distribution-Aware Reweighting
by: Paxton, Kuniko, et al.
Published: (2025)
by: Paxton, Kuniko, et al.
Published: (2025)
Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
by: Yang, Puning, et al.
Published: (2025)
by: Yang, Puning, et al.
Published: (2025)
Enhancing Adversarial Training via Reweighting Optimization Trajectory
by: Huang, Tianjin, et al.
Published: (2023)
by: Huang, Tianjin, et al.
Published: (2023)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Mitigating Mismatch within Reference-based Preference Optimization
by: Yuan, Suqin, et al.
Published: (2026)
by: Yuan, Suqin, et al.
Published: (2026)
BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning
by: Gong, Shijin, et al.
Published: (2026)
by: Gong, Shijin, et al.
Published: (2026)
Think-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded Reasoning
by: Wei, Lei, et al.
Published: (2026)
by: Wei, Lei, et al.
Published: (2026)
Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
LAD: Learning Advantage Distribution for Reasoning
by: Li, Wendi, et al.
Published: (2026)
by: Li, Wendi, et al.
Published: (2026)
Classical Verification of Quantum Learning Advantages with Noises
by: Ma, Yinghao, et al.
Published: (2024)
by: Ma, Yinghao, et al.
Published: (2024)
Mitigating the Backdoor Effect for Multi-Task Model Merging via Safety-Aware Subspace
by: Yang, Jinluan, et al.
Published: (2024)
by: Yang, Jinluan, et al.
Published: (2024)
DAJ: Data-Reweighted LLM Judge for Test-Time Scaling in Code Generation
by: Qin, Peijia, et al.
Published: (2026)
by: Qin, Peijia, et al.
Published: (2026)
Group & Reweight: A Novel Cost-Sensitive Approach to Mitigating Class Imbalance in Network Traffic Classification
by: Du, Wumei, et al.
Published: (2024)
by: Du, Wumei, et al.
Published: (2024)
DFedReweighting: A Unified Framework for Objective-Oriented Reweighting in Decentralized Federated Learning
by: Zhang, Kaichuang, et al.
Published: (2025)
by: Zhang, Kaichuang, et al.
Published: (2025)
Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning
by: Gai, Jingchu, et al.
Published: (2025)
by: Gai, Jingchu, et al.
Published: (2025)
Similar Items
-
RevOrder: A Novel Method for Enhanced Arithmetic in Language Models
by: Shen, Si, et al.
Published: (2024) -
LLM-Metrics: Measuring Research Impact Through Large Language Model Memory
by: Shen, Si, et al.
Published: (2026) -
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
by: Sun, Ke, et al.
Published: (2026) -
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
by: Li, Sijia, et al.
Published: (2026) -
TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning
by: Pan, Muyu, et al.
Published: (2026)