LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Beomseok, Song, Jiwon, Kim, Jae-Joon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918274706440192
author Kang, Beomseok
Song, Jiwon
Kim, Jae-Joon
author_facet Kang, Beomseok
Song, Jiwon
Kim, Jae-Joon
contents Multi-stage reasoning has emerged as an effective strategy for enhancing the reasoning capability of small language models by decomposing complex problems into sequential sub-stages. However, this comes at the cost of increased latency. We observe that existing adaptive acceleration techniques, such as layer skipping, struggle to balance efficiency and accuracy in this setting due to two key challenges: (1) stage-wise variation in skip sensitivity, and (2) the generation of redundant output tokens. To address these, we propose LiteStage, a latency-aware layer skipping framework for multi-stage reasoning. LiteStage combines a stage-wise offline search that allocates optimal layer budgets with an online confidence-based generation early exit to suppress unnecessary decoding. Experiments on three benchmarks, e.g., OBQA, CSQA, and StrategyQA, show that LiteStage outperforms prior training-free layer skipping methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14211
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
Kang, Beomseok
Song, Jiwon
Kim, Jae-Joon
Computation and Language
Artificial Intelligence
Multi-stage reasoning has emerged as an effective strategy for enhancing the reasoning capability of small language models by decomposing complex problems into sequential sub-stages. However, this comes at the cost of increased latency. We observe that existing adaptive acceleration techniques, such as layer skipping, struggle to balance efficiency and accuracy in this setting due to two key challenges: (1) stage-wise variation in skip sensitivity, and (2) the generation of redundant output tokens. To address these, we propose LiteStage, a latency-aware layer skipping framework for multi-stage reasoning. LiteStage combines a stage-wise offline search that allocates optimal layer budgets with an online confidence-based generation early exit to suppress unnecessary decoding. Experiments on three benchmarks, e.g., OBQA, CSQA, and StrategyQA, show that LiteStage outperforms prior training-free layer skipping methods.
title LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.14211