Test-time Prompt Intervention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chenxu, Si, Qingyi, Dai, Mz, Yao, Dingyu, Zheng, Mingyu, Chen, Minghui, Lin, Zheng, Wang, Weiping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915569917231104
author Yang, Chenxu
Si, Qingyi
Dai, Mz
Yao, Dingyu
Zheng, Mingyu
Chen, Minghui
Lin, Zheng
Wang, Weiping
author_facet Yang, Chenxu
Si, Qingyi
Dai, Mz
Yao, Dingyu
Zheng, Mingyu
Chen, Minghui
Lin, Zheng
Wang, Weiping
contents Test-time compute has led to remarkable success in the large language model (LLM) community, particularly for complex tasks, where longer chains of thought (CoTs) are generated to enhance reasoning capabilities. However, growing evidence reveals that such reasoning models often produce CoTs plagued by excessive redundancy, including unnecessary verification steps and repetitive reasoning shifts. The root cause lies in post-training of them that overly rely on outcome reward paradigms, as the data of process reward paradigms, which regulate intermediate reasoning steps, is difficult to construct at scale. To address this, we propose PI, a novel framework for Test-time Prompt Intervention. PI provides an interface to dynamically guide and regulate reasoning paths during inference through timely (When module) and proper (How module) interventions and post-intervention sampling (Which module). This allows human problem-solving expertise and cognitive science principles to be seamlessly integrated into LLMs' reasoning processes, enhancing controllability and interpretability. Extensive experiments across multiple models and datasets demonstrate that PI significantly shortens CoTs while reducing hallucination, yielding more concise and reliable reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02511
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Test-time Prompt Intervention
Yang, Chenxu
Si, Qingyi
Dai, Mz
Yao, Dingyu
Zheng, Mingyu
Chen, Minghui
Lin, Zheng
Wang, Weiping
Artificial Intelligence
Computation and Language
Test-time compute has led to remarkable success in the large language model (LLM) community, particularly for complex tasks, where longer chains of thought (CoTs) are generated to enhance reasoning capabilities. However, growing evidence reveals that such reasoning models often produce CoTs plagued by excessive redundancy, including unnecessary verification steps and repetitive reasoning shifts. The root cause lies in post-training of them that overly rely on outcome reward paradigms, as the data of process reward paradigms, which regulate intermediate reasoning steps, is difficult to construct at scale. To address this, we propose PI, a novel framework for Test-time Prompt Intervention. PI provides an interface to dynamically guide and regulate reasoning paths during inference through timely (When module) and proper (How module) interventions and post-intervention sampling (Which module). This allows human problem-solving expertise and cognitive science principles to be seamlessly integrated into LLMs' reasoning processes, enhancing controllability and interpretability. Extensive experiments across multiple models and datasets demonstrate that PI significantly shortens CoTs while reducing hallucination, yielding more concise and reliable reasoning.
title Test-time Prompt Intervention
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.02511