ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Hyunseok, Oh, Seunghyuk, Kim, Jaehyung, Shin, Jinwoo, Tack, Jihoon
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913941766012928
author Lee, Hyunseok
Oh, Seunghyuk
Kim, Jaehyung
Shin, Jinwoo
Tack, Jihoon
author_facet Lee, Hyunseok
Oh, Seunghyuk
Kim, Jaehyung
Shin, Jinwoo
Tack, Jihoon
contents Self-awareness, i.e., the ability to assess and correct one's own generation, is a fundamental aspect of human intelligence, making its replication in large language models (LLMs) an important yet challenging task. Previous works tackle this by employing extensive reinforcement learning or rather relying on large external verifiers. In this work, we propose Refine via Intrinsic Self-Verification (ReVISE), an efficient and effective framework that enables LLMs to self-correct their outputs through self-verification. The core idea of ReVISE is to enable LLMs to verify their reasoning processes and continually rethink reasoning trajectories based on its verification. We introduce a structured curriculum based upon online preference learning to implement this efficiently. Specifically, as ReVISE involves two challenging tasks (i.e., self-verification and reasoning correction), we tackle each task sequentially using curriculum learning, collecting both failed and successful reasoning paths to construct preference pairs for efficient training. During inference, our approach enjoys natural test-time scaling by integrating self-verification and correction capabilities, further enhanced by our proposed confidence-aware decoding mechanism. Our experiments on various reasoning tasks demonstrate that ReVISE achieves efficient self-correction and significantly improves reasoning performance.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14565
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification
Lee, Hyunseok
Oh, Seunghyuk
Kim, Jaehyung
Shin, Jinwoo
Tack, Jihoon
Machine Learning
Computation and Language
Self-awareness, i.e., the ability to assess and correct one's own generation, is a fundamental aspect of human intelligence, making its replication in large language models (LLMs) an important yet challenging task. Previous works tackle this by employing extensive reinforcement learning or rather relying on large external verifiers. In this work, we propose Refine via Intrinsic Self-Verification (ReVISE), an efficient and effective framework that enables LLMs to self-correct their outputs through self-verification. The core idea of ReVISE is to enable LLMs to verify their reasoning processes and continually rethink reasoning trajectories based on its verification. We introduce a structured curriculum based upon online preference learning to implement this efficiently. Specifically, as ReVISE involves two challenging tasks (i.e., self-verification and reasoning correction), we tackle each task sequentially using curriculum learning, collecting both failed and successful reasoning paths to construct preference pairs for efficient training. During inference, our approach enjoys natural test-time scaling by integrating self-verification and correction capabilities, further enhanced by our proposed confidence-aware decoding mechanism. Our experiments on various reasoning tasks demonstrate that ReVISE achieves efficient self-correction and significantly improves reasoning performance.
title ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2502.14565