CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jin, Chen, Tanno, Ryutaro, Diethe, Tom, Teare, Philip
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914315900026880
author Jin, Chen
Tanno, Ryutaro
Diethe, Tom
Teare, Philip
author_facet Jin, Chen
Tanno, Ryutaro
Diethe, Tom
Teare, Philip
contents Large Language Models (LLMs) often rely on test-time scaling via parallel decoding (for example, 512 samples) to boost reasoning accuracy, but this incurs substantial compute. We introduce CoRefine, a confidence-guided self-refinement method that achieves competitive accuracy using a fraction of the tokens via a lightweight 211k-parameter Conv1D controller atop a frozen LLM. The controller consumes full-trace confidence to decide whether to halt, re-examine, or try a different approach, enabling targeted self-correction with an average of 2.7 refinement steps per problem and roughly 190-fold token reduction relative to 512-sample baselines. Across diverse reasoning benchmarks and three open-source models, the controller achieves 92.6 percent precision when it confidently halts, indicating that confidence dynamics reliably signal correctness without ground-truth verification. We extend this to CoRefine-Tree, a hybrid sequential-parallel variant that adaptively balances exploration and exploitation, with easy serving integration and verifier compatibility. By treating confidence as a control signal rather than a correctness guarantee, CoRefine provides a modular primitive for scalable reasoning and agentic settings with imperfect verifiers.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08948
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute
Jin, Chen
Tanno, Ryutaro
Diethe, Tom
Teare, Philip
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) often rely on test-time scaling via parallel decoding (for example, 512 samples) to boost reasoning accuracy, but this incurs substantial compute. We introduce CoRefine, a confidence-guided self-refinement method that achieves competitive accuracy using a fraction of the tokens via a lightweight 211k-parameter Conv1D controller atop a frozen LLM. The controller consumes full-trace confidence to decide whether to halt, re-examine, or try a different approach, enabling targeted self-correction with an average of 2.7 refinement steps per problem and roughly 190-fold token reduction relative to 512-sample baselines. Across diverse reasoning benchmarks and three open-source models, the controller achieves 92.6 percent precision when it confidently halts, indicating that confidence dynamics reliably signal correctness without ground-truth verification. We extend this to CoRefine-Tree, a hybrid sequential-parallel variant that adaptively balances exploration and exploitation, with easy serving integration and verifier compatibility. By treating confidence as a control signal rather than a correctness guarantee, CoRefine provides a modular primitive for scalable reasoning and agentic settings with imperfect verifiers.
title CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.08948