Efficient Thought Space Exploration Through Strategic Intervention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Ziheng, Cai, Hengyi, Wei, Xiaochi, Li, Yuchen, Wang, Shuaiqiang, Deng, Zhi-Hong, Yin, Dawei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912865497120768
author Li, Ziheng
Cai, Hengyi
Wei, Xiaochi
Li, Yuchen
Wang, Shuaiqiang
Deng, Zhi-Hong
Yin, Dawei
author_facet Li, Ziheng
Cai, Hengyi
Wei, Xiaochi
Li, Yuchen
Wang, Shuaiqiang
Deng, Zhi-Hong
Yin, Dawei
contents While large language models (LLMs) demonstrate emerging reasoning capabilities, current inference-time expansion methods incur prohibitive computational costs by exhaustive sampling. Through analyzing decoding trajectories, we observe that most next-token predictions align well with the golden output, except for a few critical tokens that lead to deviations. Inspired by this phenomenon, we propose a novel Hint-Practice Reasoning (HPR) framework that operationalizes this insight through two synergistic components: 1) a hinter (powerful LLM) that provides probabilistic guidance at critical decision points, and 2) a practitioner (efficient smaller model) that executes major reasoning steps. The framework's core innovation lies in Distributional Inconsistency Reduction (DIR), a theoretically-grounded metric that dynamically identifies intervention points by quantifying the divergence between practitioner's reasoning trajectory and hinter's expected distribution in a tree-structured probabilistic space. Through iterative tree updates guided by DIR, HPR reweights promising reasoning paths while deprioritizing low-probability branches. Experiments across arithmetic and commonsense reasoning benchmarks demonstrate HPR's state-of-the-art efficiency-accuracy tradeoffs: it achieves comparable performance to self-consistency and MCTS baselines while decoding only 1/5 tokens, and outperforms existing methods by at most 5.1% absolute accuracy while maintaining similar or lower FLOPs.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Thought Space Exploration Through Strategic Intervention
Li, Ziheng
Cai, Hengyi
Wei, Xiaochi
Li, Yuchen
Wang, Shuaiqiang
Deng, Zhi-Hong
Yin, Dawei
Artificial Intelligence
While large language models (LLMs) demonstrate emerging reasoning capabilities, current inference-time expansion methods incur prohibitive computational costs by exhaustive sampling. Through analyzing decoding trajectories, we observe that most next-token predictions align well with the golden output, except for a few critical tokens that lead to deviations. Inspired by this phenomenon, we propose a novel Hint-Practice Reasoning (HPR) framework that operationalizes this insight through two synergistic components: 1) a hinter (powerful LLM) that provides probabilistic guidance at critical decision points, and 2) a practitioner (efficient smaller model) that executes major reasoning steps. The framework's core innovation lies in Distributional Inconsistency Reduction (DIR), a theoretically-grounded metric that dynamically identifies intervention points by quantifying the divergence between practitioner's reasoning trajectory and hinter's expected distribution in a tree-structured probabilistic space. Through iterative tree updates guided by DIR, HPR reweights promising reasoning paths while deprioritizing low-probability branches. Experiments across arithmetic and commonsense reasoning benchmarks demonstrate HPR's state-of-the-art efficiency-accuracy tradeoffs: it achieves comparable performance to self-consistency and MCTS baselines while decoding only 1/5 tokens, and outperforms existing methods by at most 5.1% absolute accuracy while maintaining similar or lower FLOPs.
title Efficient Thought Space Exploration Through Strategic Intervention
topic Artificial Intelligence
url https://arxiv.org/abs/2511.10038