Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yunqiao, Ren, Houxing, Lu, Zimu, Wang, Ke, Shi, Weikang, Zhou, Aojun, Pan, Junting, Zhan, Mingjie, Li, Hongsheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908384288047104
author Yang, Yunqiao
Ren, Houxing
Lu, Zimu
Wang, Ke
Shi, Weikang
Zhou, Aojun
Pan, Junting
Zhan, Mingjie
Li, Hongsheng
author_facet Yang, Yunqiao
Ren, Houxing
Lu, Zimu
Wang, Ke
Shi, Weikang
Zhou, Aojun
Pan, Junting
Zhan, Mingjie
Li, Hongsheng
contents Recent advances in preference optimization have demonstrated significant potential for improving mathematical reasoning capabilities in large language models (LLMs). While current approaches leverage high-quality pairwise preference data through outcome-based criteria like answer correctness or consistency, they fundamentally neglect the internal logical coherence of responses. To overcome this, we propose Probability-Consistent Preference Optimization (PCPO), a novel framework that establishes dual quantitative metrics for preference selection: (1) surface-level answer correctness and (2) intrinsic token-level probability consistency across responses. Extensive experiments show that our PCPO consistently outperforms existing outcome-only criterion approaches across a diverse range of LLMs and benchmarks. Our code is publicly available at https://github.com/YunqiaoYang/PCPO.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23540
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
Yang, Yunqiao
Ren, Houxing
Lu, Zimu
Wang, Ke
Shi, Weikang
Zhou, Aojun
Pan, Junting
Zhan, Mingjie
Li, Hongsheng
Computation and Language
Recent advances in preference optimization have demonstrated significant potential for improving mathematical reasoning capabilities in large language models (LLMs). While current approaches leverage high-quality pairwise preference data through outcome-based criteria like answer correctness or consistency, they fundamentally neglect the internal logical coherence of responses. To overcome this, we propose Probability-Consistent Preference Optimization (PCPO), a novel framework that establishes dual quantitative metrics for preference selection: (1) surface-level answer correctness and (2) intrinsic token-level probability consistency across responses. Extensive experiments show that our PCPO consistently outperforms existing outcome-only criterion approaches across a diverse range of LLMs and benchmarks. Our code is publicly available at https://github.com/YunqiaoYang/PCPO.
title Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
topic Computation and Language
url https://arxiv.org/abs/2505.23540