Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Changyuan, Lu, Zhicong, Qian, Shuang, Liu, Nayu, Li, Peiguang, Jin, Li, Hu, Leiyi, Zeng, Zhizhao, Wang, Sirui, Zeng, Ke, Guo, Zhi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
by: Lu, Zhicong, et al.
Published: (2026)
by: Lu, Zhicong, et al.
Published: (2026)
Feature-Aware Malicious Output Detection and Mitigation
by: Dong, Weilong, et al.
Published: (2025)
by: Dong, Weilong, et al.
Published: (2025)
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
LLMs Know What They Need: Leveraging a Missing Information Guided Framework to Empower Retrieval-Augmented Generation
by: Wang, Keheng, et al.
Published: (2024)
by: Wang, Keheng, et al.
Published: (2024)
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy
by: Zhang, Xiaoyun, et al.
Published: (2025)
by: Zhang, Xiaoyun, et al.
Published: (2025)
Beyond the Sequence: Statistics-Driven Pre-training for Stabilizing Sequential Recommendation Model
by: Wang, Sirui, et al.
Published: (2024)
by: Wang, Sirui, et al.
Published: (2024)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
by: Tian, Changyuan, et al.
Published: (2026)
by: Tian, Changyuan, et al.
Published: (2026)
MARGE: Improving Math Reasoning for LLMs with Guided Exploration
by: Gao, Jingyue, et al.
Published: (2025)
by: Gao, Jingyue, et al.
Published: (2025)
WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning
by: Zhuang, Wenwen, et al.
Published: (2024)
by: Zhuang, Wenwen, et al.
Published: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs
by: Cheng, Letian, et al.
Published: (2026)
by: Cheng, Letian, et al.
Published: (2026)
Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions
by: Kou, Siqi, et al.
Published: (2025)
by: Kou, Siqi, et al.
Published: (2025)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
by: Tian, Shi-Yu, et al.
Published: (2025)
by: Tian, Shi-Yu, et al.
Published: (2025)
The Perplexity Paradox: Why Code Compresses Better Than Math in LLM Prompts
by: Johnson, Warren
Published: (2026)
by: Johnson, Warren
Published: (2026)
Abstract of Structural Critique of Reasoning
by: Zixi, Li
Published: (2025)
by: Zixi, Li
Published: (2025)
Arrows of Math Reasoning Data Synthesis for Large Language Models: Diversity, Complexity and Correctness
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
by: Cui, Yingqian, et al.
Published: (2025)
by: Cui, Yingqian, et al.
Published: (2025)
Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation
by: Wang, Junjie, et al.
Published: (2026)
by: Wang, Junjie, et al.
Published: (2026)
HyCoRA: Hyper-Contrastive Role-Adaptive Learning for Role-Playing
by: Yang, Shihao, et al.
Published: (2025)
by: Yang, Shihao, et al.
Published: (2025)
Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMs
by: Mukherjee, Sagnik, et al.
Published: (2025)
by: Mukherjee, Sagnik, et al.
Published: (2025)
Physics-Guided Rectified Flow for Low-light RAW Image Enhancement
by: Zeng, Juntai
Published: (2025)
by: Zeng, Juntai
Published: (2025)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Neuro-Symbolic Data Generation for Math Reasoning
by: Li, Zenan, et al.
Published: (2024)
by: Li, Zenan, et al.
Published: (2024)
Debiasing Sequential Recommendation with Time-aware Inverse Propensity Scoring
by: Huang, Sirui, et al.
Published: (2026)
by: Huang, Sirui, et al.
Published: (2026)
Global strong solutions to the frame hydrodynamics for biaxial nematic phases
by: Feng, Minjiang, et al.
Published: (2025)
by: Feng, Minjiang, et al.
Published: (2025)
Rigorous uniaxial limit of the Qian--Sheng inertial Q-tensor hydrodynamics for liquid crystals
by: Li, Sirui, et al.
Published: (2024)
by: Li, Sirui, et al.
Published: (2024)
PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
by: Guan, Xinyu, et al.
Published: (2025)
by: Guan, Xinyu, et al.
Published: (2025)
Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique
by: Li, Yansi, et al.
Published: (2025)
by: Li, Yansi, et al.
Published: (2025)
From Good to Great: Improving Math Reasoning with Tool-Augmented Interleaf Prompting
by: Chen, Nuo, et al.
Published: (2023)
by: Chen, Nuo, et al.
Published: (2023)
Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning
by: Zeng, Fanhu, et al.
Published: (2026)
by: Zeng, Fanhu, et al.
Published: (2026)
Poisoning mechanism of ammonia on proton transport and ionomer structure in cathode catalyst layer of PEM fuel cells
by: Huang, Yichao, et al.
Published: (2026)
by: Huang, Yichao, et al.
Published: (2026)
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning
by: Hu, Wenbin, et al.
Published: (2025)
by: Hu, Wenbin, et al.
Published: (2025)
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
by: Ankner, Zachary, et al.
Published: (2024)
by: Ankner, Zachary, et al.
Published: (2024)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
by: Huang, Kaixuan, et al.
Published: (2025)
by: Huang, Kaixuan, et al.
Published: (2025)
SOLIDGEO: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
by: Wang, Peijie, et al.
Published: (2025)
by: Wang, Peijie, et al.
Published: (2025)
MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs
by: Pati, Viresh, et al.
Published: (2026)
by: Pati, Viresh, et al.
Published: (2026)
ADL: A Declarative Language for Agent-Based Chatbots
by: Zeng, Sirui, et al.
Published: (2025)
by: Zeng, Sirui, et al.
Published: (2025)
Similar Items
-
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
by: Lu, Zhicong, et al.
Published: (2026) -
Feature-Aware Malicious Output Detection and Mitigation
by: Dong, Weilong, et al.
Published: (2025) -
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
by: Xu, Yifan, et al.
Published: (2024) -
LLMs Know What They Need: Leveraging a Missing Information Guided Framework to Empower Retrieval-Augmented Generation
by: Wang, Keheng, et al.
Published: (2024) -
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy
by: Zhang, Xiaoyun, et al.
Published: (2025)