Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hong, Colin, Guo, Xu, Singh, Anand Chaanan, Choukse, Esha, Ustiugov, Dmitrii
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918143035703296
author Hong, Colin
Guo, Xu
Singh, Anand Chaanan
Choukse, Esha
Ustiugov, Dmitrii
author_facet Hong, Colin
Guo, Xu
Singh, Anand Chaanan
Choukse, Esha
Ustiugov, Dmitrii
contents Recently, Test-Time Scaling (TTS) has gained increasing attention for improving LLM reasoning performance at test time without retraining the model. A notable TTS technique is Self-Consistency (SC), which generates multiple reasoning chains in parallel and selects the final answer via majority voting. While effective, the order-of-magnitude computational overhead limits its broad deployment. Prior attempts to accelerate SC mainly rely on model-based confidence scores or heuristics with limited empirical support. For the first time, we theoretically and empirically analyze the inefficiencies of SC and reveal actionable opportunities for improvement. Building on these insights, we propose Slim-SC, a step-wise pruning strategy that identifies and removes redundant chains using inter-chain similarity at the thought level. Experiments on three STEM reasoning datasets and two recent LLM architectures show that Slim-SC reduces inference latency and KVC usage by up to 45% and 26%, respectively, with R1-Distill, while maintaining or improving accuracy, thus offering a simple yet efficient TTS alternative for SC.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13990
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency
Hong, Colin
Guo, Xu
Singh, Anand Chaanan
Choukse, Esha
Ustiugov, Dmitrii
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
Recently, Test-Time Scaling (TTS) has gained increasing attention for improving LLM reasoning performance at test time without retraining the model. A notable TTS technique is Self-Consistency (SC), which generates multiple reasoning chains in parallel and selects the final answer via majority voting. While effective, the order-of-magnitude computational overhead limits its broad deployment. Prior attempts to accelerate SC mainly rely on model-based confidence scores or heuristics with limited empirical support. For the first time, we theoretically and empirically analyze the inefficiencies of SC and reveal actionable opportunities for improvement. Building on these insights, we propose Slim-SC, a step-wise pruning strategy that identifies and removes redundant chains using inter-chain similarity at the thought level. Experiments on three STEM reasoning datasets and two recent LLM architectures show that Slim-SC reduces inference latency and KVC usage by up to 45% and 26%, respectively, with R1-Distill, while maintaining or improving accuracy, thus offering a simple yet efficient TTS alternative for SC.
title Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7
url https://arxiv.org/abs/2509.13990