Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yunqiao, Ren, Houxing, Lu, Zimu, Wang, Ke, Shi, Weikang, Zhou, Aojun, Pan, Junting, Zhan, Mingjie, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Alignment with Fill-In-the-Middle for Enhancing Code Generation
by: Ren, Houxing, et al.
Published: (2025)
by: Ren, Houxing, et al.
Published: (2025)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
by: Shi, Weikang, et al.
Published: (2026)
by: Shi, Weikang, et al.
Published: (2026)
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
Edit-Based Refinement for Parallel Masked Diffusion Language Models
by: Ren, Houxing, et al.
Published: (2026)
by: Ren, Houxing, et al.
Published: (2026)
Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
by: Ren, Houxing, et al.
Published: (2026)
by: Ren, Houxing, et al.
Published: (2026)
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
by: Ren, Houxing, et al.
Published: (2024)
by: Ren, Houxing, et al.
Published: (2024)
FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
by: Lu, Zimu, et al.
Published: (2026)
by: Lu, Zimu, et al.
Published: (2026)
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
by: Wang, Ke, et al.
Published: (2024)
by: Wang, Ke, et al.
Published: (2024)
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics
by: Yang, Yunqiao, et al.
Published: (2026)
by: Yang, Yunqiao, et al.
Published: (2026)
Empowering Character-level Text Infilling by Eliminating Sub-Tokens
by: Ren, Houxing, et al.
Published: (2024)
by: Ren, Houxing, et al.
Published: (2024)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
by: Shi, Weikang, et al.
Published: (2025)
by: Shi, Weikang, et al.
Published: (2025)
LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
by: Lu, Junjie, et al.
Published: (2025)
by: Lu, Junjie, et al.
Published: (2025)
Genetic Auto-prompt Learning for Pre-trained Code Intelligence Language Models
by: Feng, Chengzhe, et al.
Published: (2024)
by: Feng, Chengzhe, et al.
Published: (2024)
Topology-Enhanced Alignment for Large Language Models: Trajectory Topology Loss and Topological Preference Optimization
by: Pan, Yurui, et al.
Published: (2026)
by: Pan, Yurui, et al.
Published: (2026)
MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control
by: Wang, Yuchi, et al.
Published: (2026)
by: Wang, Yuchi, et al.
Published: (2026)
Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning
by: Wang, Tianduo, et al.
Published: (2024)
by: Wang, Tianduo, et al.
Published: (2024)
Do LLM Evaluators Prefer Themselves for a Reason?
by: Chen, Wei-Lin, et al.
Published: (2025)
by: Chen, Wei-Lin, et al.
Published: (2025)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
by: Zhang, Renrui, et al.
Published: (2024)
by: Zhang, Renrui, et al.
Published: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Self-Consistency Preference Optimization
by: Prasad, Archiki, et al.
Published: (2024)
by: Prasad, Archiki, et al.
Published: (2024)
DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization
by: Deng, Mengyi, et al.
Published: (2026)
by: Deng, Mengyi, et al.
Published: (2026)
NODI: Out-Of-Distribution Detection with Noise from Diffusion
by: Zhou, Jingqiu, et al.
Published: (2024)
by: Zhou, Jingqiu, et al.
Published: (2024)
RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
by: Hsu, Chia-Hsuan, et al.
Published: (2025)
by: Hsu, Chia-Hsuan, et al.
Published: (2025)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2025)
by: Xiao, Han, et al.
Published: (2025)
Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game
by: Cheng, Pengyu, et al.
Published: (2023)
by: Cheng, Pengyu, et al.
Published: (2023)
ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
by: Pan, Haolin, et al.
Published: (2026)
by: Pan, Haolin, et al.
Published: (2026)
Enhancing LLM Safety via Constrained Direct Preference Optimization
by: Liu, Zixuan, et al.
Published: (2024)
by: Liu, Zixuan, et al.
Published: (2024)
CSCE: Boosting LLM Reasoning by Simultaneous Enhancing of Causal Significance and Consistency
by: Wang, Kangsheng, et al.
Published: (2024)
by: Wang, Kangsheng, et al.
Published: (2024)
HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression
by: Zheng, Minghui, et al.
Published: (2026)
by: Zheng, Minghui, et al.
Published: (2026)
Budget-Aware Anytime Reasoning with LLM-Synthesized Preference Data
by: Zhang, Xuanming, et al.
Published: (2026)
by: Zhang, Xuanming, et al.
Published: (2026)
Preference Optimization for Reasoning with Pseudo Feedback
by: Jiao, Fangkai, et al.
Published: (2024)
by: Jiao, Fangkai, et al.
Published: (2024)
Similar Items
-
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning
by: Lu, Zimu, et al.
Published: (2024) -
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
by: Lu, Zimu, et al.
Published: (2024) -
Alignment with Fill-In-the-Middle for Enhancing Code Generation
by: Ren, Houxing, et al.
Published: (2025) -
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
by: Lu, Zimu, et al.
Published: (2024) -
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
by: Wang, Ke, et al.
Published: (2025)