Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Long, Quanyu, Jiang, Kai Jie, Chen, Jianda, Guo, Xu, Gan, Leilei, Wang, Wenya
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914304123469824
author Long, Quanyu
Jiang, Kai Jie
Chen, Jianda
Guo, Xu
Gan, Leilei
Wang, Wenya
author_facet Long, Quanyu
Jiang, Kai Jie
Chen, Jianda
Guo, Xu
Gan, Leilei
Wang, Wenya
contents Large Reasoning Models (LRMs) achieve strong performance by generating long reasoning traces with reflection. Through a large-scale empirical analysis, we find that a substantial fraction of reflective steps consist of self-verification (recheck) that repeatedly confirm intermediate results. These rechecks occur frequently across models and benchmarks, yet the vast majority are confirmatory rather than corrective, rarely identifying errors and altering reasoning outcomes. This reveals a mismatch between how often self-verification is activated and how often it is actually useful. Motivated by this, we propose a novel, experience-driven test-time framework that reduces the overused verification. Our method detects the activation of recheck behavior, consults an offline experience pool of past verification outcomes, and estimates whether a recheck is likely unnecessary via efficient retrieval. When historical experience suggests unnecessary, a suppression signal redirects the model to proceed. Across multiple model and benchmarks, our approach reduces token usage up to 20.3% while maintaining the accuracy, and in some datasets even yields accuracy improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03485
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning
Long, Quanyu
Jiang, Kai Jie
Chen, Jianda
Guo, Xu
Gan, Leilei
Wang, Wenya
Computation and Language
Artificial Intelligence
Machine Learning
Large Reasoning Models (LRMs) achieve strong performance by generating long reasoning traces with reflection. Through a large-scale empirical analysis, we find that a substantial fraction of reflective steps consist of self-verification (recheck) that repeatedly confirm intermediate results. These rechecks occur frequently across models and benchmarks, yet the vast majority are confirmatory rather than corrective, rarely identifying errors and altering reasoning outcomes. This reveals a mismatch between how often self-verification is activated and how often it is actually useful. Motivated by this, we propose a novel, experience-driven test-time framework that reduces the overused verification. Our method detects the activation of recheck behavior, consults an offline experience pool of past verification outcomes, and estimates whether a recheck is likely unnecessary via efficient retrieval. When historical experience suggests unnecessary, a suppression signal redirects the model to proceed. Across multiple model and benchmarks, our approach reduces token usage up to 20.3% while maintaining the accuracy, and in some datasets even yields accuracy improvements.
title Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.03485