Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhong, Jianyuan, Li, Zeju, Xu, Zhijian, Wen, Xiangyu, Li, Kezhi, Xu, Qiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913844711915520
author Zhong, Jianyuan
Li, Zeju
Xu, Zhijian
Wen, Xiangyu
Li, Kezhi
Xu, Qiang
author_facet Zhong, Jianyuan
Li, Zeju
Xu, Zhijian
Wen, Xiangyu
Li, Kezhi
Xu, Qiang
contents Large Language Model (LLM) reasoning for complex tasks inherently involves a trade-off between solution accuracy and computational efficiency. The subsequent step of verification, while intended to improve performance, further complicates this landscape by introducing its own challenging trade-off: sophisticated Generative Reward Models (GenRMs) can be computationally prohibitive if naively integrated with LLMs at test-time, while simpler, faster methods may lack reliability. To overcome these challenges, we introduce FlexiVe, a novel generative verifier that flexibly balances computational resources between rapid, reliable fast thinking and meticulous slow thinking using a Flexible Allocation of Verification Budget strategy. We further propose the Solve-Detect-Verify pipeline, an efficient inference-time scaling framework that intelligently integrates FlexiVe, proactively identifying solution completion points to trigger targeted verification and provide focused solver feedback. Experiments show FlexiVe achieves superior accuracy in pinpointing errors within reasoning traces on ProcessBench. Furthermore, on challenging mathematical reasoning benchmarks (AIME 2024, AIME 2025, and CNMO), our full approach outperforms baselines like self-consistency in reasoning accuracy and inference efficiency. Our system offers a scalable and effective solution to enhance LLM reasoning at test time.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11966
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
Zhong, Jianyuan
Li, Zeju
Xu, Zhijian
Wen, Xiangyu
Li, Kezhi
Xu, Qiang
Artificial Intelligence
Large Language Model (LLM) reasoning for complex tasks inherently involves a trade-off between solution accuracy and computational efficiency. The subsequent step of verification, while intended to improve performance, further complicates this landscape by introducing its own challenging trade-off: sophisticated Generative Reward Models (GenRMs) can be computationally prohibitive if naively integrated with LLMs at test-time, while simpler, faster methods may lack reliability. To overcome these challenges, we introduce FlexiVe, a novel generative verifier that flexibly balances computational resources between rapid, reliable fast thinking and meticulous slow thinking using a Flexible Allocation of Verification Budget strategy. We further propose the Solve-Detect-Verify pipeline, an efficient inference-time scaling framework that intelligently integrates FlexiVe, proactively identifying solution completion points to trigger targeted verification and provide focused solver feedback. Experiments show FlexiVe achieves superior accuracy in pinpointing errors within reasoning traces on ProcessBench. Furthermore, on challenging mathematical reasoning benchmarks (AIME 2024, AIME 2025, and CNMO), our full approach outperforms baselines like self-consistency in reasoning accuracy and inference efficiency. Our system offers a scalable and effective solution to enhance LLM reasoning at test time.
title Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
topic Artificial Intelligence
url https://arxiv.org/abs/2505.11966