PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Leang, Joshua Ong Jun, Zhao, Zheng, Gema, Aryo Pradipta, Yang, Sohee, Kwan, Wai-Chung, He, Xuanli, Li, Wenda, Minervini, Pasquale, Giunchiglia, Eleonora, Cohen, Shay B.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918474697146368
author Leang, Joshua Ong Jun
Zhao, Zheng
Gema, Aryo Pradipta
Yang, Sohee
Kwan, Wai-Chung
He, Xuanli
Li, Wenda
Minervini, Pasquale
Giunchiglia, Eleonora
Cohen, Shay B.
author_facet Leang, Joshua Ong Jun
Zhao, Zheng
Gema, Aryo Pradipta
Yang, Sohee
Kwan, Wai-Chung
He, Xuanli
Li, Wenda
Minervini, Pasquale
Giunchiglia, Eleonora
Cohen, Shay B.
contents Best-of-n sampling improves the accuracy of large language models (LLMs) and large reasoning models (LRMs) by generating multiple candidate solutions and selecting the one with the highest reward. The key challenge for reasoning tasks is designing a scoring function that can identify correct reasoning chains without access to ground-truth answers. We propose Probabilistic Confidence Selection And Ranking (PiCSAR): a simple, training-free method that scores each candidate generation using the joint log-likelihood of the reasoning and final answer. The joint log-likelihood of the reasoning and final answer naturally decomposes into reasoning confidence and answer confidence. PiCSAR achieves substantial gains across diverse benchmarks (+10.18 on MATH500, +9.81 on AIME2025), outperforming baselines with at least 2x fewer samples in 16 out of 20 comparisons. Our analysis reveals that correct reasoning chains exhibit significantly higher reasoning and answer confidence, justifying the effectiveness of PiCSAR.
format Preprint
id arxiv_https___arxiv_org_abs_2508_21787
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
Leang, Joshua Ong Jun
Zhao, Zheng
Gema, Aryo Pradipta
Yang, Sohee
Kwan, Wai-Chung
He, Xuanli
Li, Wenda
Minervini, Pasquale
Giunchiglia, Eleonora
Cohen, Shay B.
Computation and Language
Artificial Intelligence
Best-of-n sampling improves the accuracy of large language models (LLMs) and large reasoning models (LRMs) by generating multiple candidate solutions and selecting the one with the highest reward. The key challenge for reasoning tasks is designing a scoring function that can identify correct reasoning chains without access to ground-truth answers. We propose Probabilistic Confidence Selection And Ranking (PiCSAR): a simple, training-free method that scores each candidate generation using the joint log-likelihood of the reasoning and final answer. The joint log-likelihood of the reasoning and final answer naturally decomposes into reasoning confidence and answer confidence. PiCSAR achieves substantial gains across diverse benchmarks (+10.18 on MATH500, +9.81 on AIME2025), outperforming baselines with at least 2x fewer samples in 16 out of 20 comparisons. Our analysis reveals that correct reasoning chains exhibit significantly higher reasoning and answer confidence, justifying the effectiveness of PiCSAR.
title PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.21787