Temporal Consistency for LLM Reasoning Process Error Identification

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Guo, Jiacheng, Wu, Yue, Qiu, Jiahao, Huang, Kaixuan, Juan, Xinzhe, Yang, Ling, Wang, Mengdi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914172093071360
author Guo, Jiacheng
Wu, Yue
Qiu, Jiahao
Huang, Kaixuan
Juan, Xinzhe
Yang, Ling
Wang, Mengdi
author_facet Guo, Jiacheng
Wu, Yue
Qiu, Jiahao
Huang, Kaixuan
Juan, Xinzhe
Yang, Ling
Wang, Mengdi
contents Verification is crucial for effective mathematical reasoning. We present a new temporal consistency method where verifiers iteratively refine their judgments based on the previous assessment. Unlike one-round verification or multi-model debate approaches, our method leverages consistency in a sequence of self-reflection actions to improve verification accuracy. Empirical evaluations across diverse mathematical process error identification benchmarks (Mathcheck, ProcessBench, and PRM800K) show consistent performance improvements over baseline methods. When applied to the recent DeepSeek R1 distilled models, our method demonstrates strong performance, enabling 7B/8B distilled models to outperform all 70B/72B models and GPT-4o on ProcessBench. Notably, the distilled 14B model with our method achieves performance comparable to Deepseek-R1. Our codes are available at https://github.com/jcguo123/Temporal-Consistency
format Preprint
id arxiv_https___arxiv_org_abs_2503_14495
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Temporal Consistency for LLM Reasoning Process Error Identification
Guo, Jiacheng
Wu, Yue
Qiu, Jiahao
Huang, Kaixuan
Juan, Xinzhe
Yang, Ling
Wang, Mengdi
Computation and Language
Artificial Intelligence
Machine Learning
Verification is crucial for effective mathematical reasoning. We present a new temporal consistency method where verifiers iteratively refine their judgments based on the previous assessment. Unlike one-round verification or multi-model debate approaches, our method leverages consistency in a sequence of self-reflection actions to improve verification accuracy. Empirical evaluations across diverse mathematical process error identification benchmarks (Mathcheck, ProcessBench, and PRM800K) show consistent performance improvements over baseline methods. When applied to the recent DeepSeek R1 distilled models, our method demonstrates strong performance, enabling 7B/8B distilled models to outperform all 70B/72B models and GPT-4o on ProcessBench. Notably, the distilled 14B model with our method achieves performance comparable to Deepseek-R1. Our codes are available at https://github.com/jcguo123/Temporal-Consistency
title Temporal Consistency for LLM Reasoning Process Error Identification
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.14495