Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Rongman, Li, Yifei, Zhao, Tianzhe, Wu, Yanrui, Li, Bo, Yan, Hang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916013386235904
author Xu, Rongman
Li, Yifei
Zhao, Tianzhe
Wu, Yanrui
Li, Bo
Yan, Hang
author_facet Xu, Rongman
Li, Yifei
Zhao, Tianzhe
Wu, Yanrui
Li, Bo
Yan, Hang
contents Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning. However, maximizing their potential through inference-time scaling faces challenges in trade-off between sampling budget and reasoning quality. Current strategies remain inefficient as they typically treat sampling width and depth as orthogonal objectives, where width consensus methods risk reinforcing hallucinations, while depth pruning mechanisms prematurely truncate complex yet valid reasoning chains. Therefore, we propose Dual-Dimensional Consistency (DDC), a unified framework that bridges path quality with adaptive termination. By coupling Confidence-Weighted Bayesian protocol with a Trend-Aware Stratified Pruning, our method ensures that computational resources are concentrated on high quality reasoning paths, filtering hallucinations while accelerating consensus. Evaluations across five benchmarks demonstrate that this approach reduces token consumption by over 10 times while maintaining or exceeding the accuracy of strong baselines across various LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15100
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling
Xu, Rongman
Li, Yifei
Zhao, Tianzhe
Wu, Yanrui
Li, Bo
Yan, Hang
Artificial Intelligence
Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning. However, maximizing their potential through inference-time scaling faces challenges in trade-off between sampling budget and reasoning quality. Current strategies remain inefficient as they typically treat sampling width and depth as orthogonal objectives, where width consensus methods risk reinforcing hallucinations, while depth pruning mechanisms prematurely truncate complex yet valid reasoning chains. Therefore, we propose Dual-Dimensional Consistency (DDC), a unified framework that bridges path quality with adaptive termination. By coupling Confidence-Weighted Bayesian protocol with a Trend-Aware Stratified Pruning, our method ensures that computational resources are concentrated on high quality reasoning paths, filtering hallucinations while accelerating consensus. Evaluations across five benchmarks demonstrate that this approach reduces token consumption by over 10 times while maintaining or exceeding the accuracy of strong baselines across various LLMs.
title Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling
topic Artificial Intelligence
url https://arxiv.org/abs/2605.15100