AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Zhuomin, Yao, Yizhen, Zuo, Pengfei, Gao, Bin, Li, Qinya, Zheng, Zhenzhe, Wu, Fan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917884121317376
author He, Zhuomin
Yao, Yizhen
Zuo, Pengfei
Gao, Bin
Li, Qinya
Zheng, Zhenzhe
Wu, Fan
author_facet He, Zhuomin
Yao, Yizhen
Zuo, Pengfei
Gao, Bin
Li, Qinya
Zheng, Zhenzhe
Wu, Fan
contents Long-context large language models (LLMs) inference is increasingly critical, motivating a number of studies devoted to alleviating the substantial storage and computational costs in such scenarios. Layer-wise skipping methods are promising optimizations but rarely explored in long-context inference. We observe that existing layer-wise skipping strategies have several limitations when applied in long-context inference, including the inability to adapt to model and context variability, disregard for sublayer significance, and inapplicability for the prefilling phase. This paper proposes \sysname, an adaptive sublayer skipping method specifically designed for long-context inference. \sysname adaptively identifies less important layers by leveraging on-the-fly similarity information, enables sublayer-wise skipping, and accelerates both the prefilling and decoding phases. The effectiveness of \sysname is demonstrated through extensive experiments on various long-context benchmarks and models, showcasing its superior inference performance over existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02336
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
He, Zhuomin
Yao, Yizhen
Zuo, Pengfei
Gao, Bin
Li, Qinya
Zheng, Zhenzhe
Wu, Fan
Computation and Language
Artificial Intelligence
Long-context large language models (LLMs) inference is increasingly critical, motivating a number of studies devoted to alleviating the substantial storage and computational costs in such scenarios. Layer-wise skipping methods are promising optimizations but rarely explored in long-context inference. We observe that existing layer-wise skipping strategies have several limitations when applied in long-context inference, including the inability to adapt to model and context variability, disregard for sublayer significance, and inapplicability for the prefilling phase. This paper proposes \sysname, an adaptive sublayer skipping method specifically designed for long-context inference. \sysname adaptively identifies less important layers by leveraging on-the-fly similarity information, enables sublayer-wise skipping, and accelerates both the prefilling and decoding phases. The effectiveness of \sysname is demonstrated through extensive experiments on various long-context benchmarks and models, showcasing its superior inference performance over existing baselines.
title AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.02336