The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jindal, Ishan, Akuthota, Sai Prashanth, Taneja, Jayant, Sharma, Sachin Dev
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912873038479360
author Jindal, Ishan
Akuthota, Sai Prashanth
Taneja, Jayant
Sharma, Sachin Dev
author_facet Jindal, Ishan
Akuthota, Sai Prashanth
Taneja, Jayant
Sharma, Sachin Dev
contents Large language models achieve strong reasoning performance, but inference strategies such as Self-Consistency (SC) are computationally expensive, as they fully expand all reasoning traces. We introduce PoLR (Path of Least Resistance), the first inference-time method to leverage prefix consistency for compute-efficient reasoning. PoLR clusters short prefixes of reasoning traces, identifies the dominant cluster, and expands all paths in that cluster, preserving the accuracy benefits of SC while substantially reducing token usage and latency. Our theoretical analysis, framed via mutual information and entropy, explains why early reasoning steps encode strong signals predictive of final correctness. Empirically, PoLR consistently matches or exceeds SC across GSM8K, MATH500, AIME24/25, and GPQA-DIAMOND, reducing token usage by up to 60% and wall-clock latency by up to 50%. Moreover, PoLR is fully complementary to adaptive inference methods (e.g., Adaptive Consistency, Early-Stopping SC) and can serve as a drop-in pre-filter, making SC substantially more efficient and scalable without requiring model fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21494
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus
Jindal, Ishan
Akuthota, Sai Prashanth
Taneja, Jayant
Sharma, Sachin Dev
Artificial Intelligence
Computation and Language
Large language models achieve strong reasoning performance, but inference strategies such as Self-Consistency (SC) are computationally expensive, as they fully expand all reasoning traces. We introduce PoLR (Path of Least Resistance), the first inference-time method to leverage prefix consistency for compute-efficient reasoning. PoLR clusters short prefixes of reasoning traces, identifies the dominant cluster, and expands all paths in that cluster, preserving the accuracy benefits of SC while substantially reducing token usage and latency. Our theoretical analysis, framed via mutual information and entropy, explains why early reasoning steps encode strong signals predictive of final correctness. Empirically, PoLR consistently matches or exceeds SC across GSM8K, MATH500, AIME24/25, and GPQA-DIAMOND, reducing token usage by up to 60% and wall-clock latency by up to 50%. Moreover, PoLR is fully complementary to adaptive inference methods (e.g., Adaptive Consistency, Early-Stopping SC) and can serve as a drop-in pre-filter, making SC substantially more efficient and scalable without requiring model fine-tuning.
title The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.21494