DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Ning, Liu, Fangxin, Wang, Junjie, Yang, Tao, Liu, Kan, Guan, Haibing, Jiang, Li
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918031934881792
author Yang, Ning
Liu, Fangxin
Wang, Junjie
Yang, Tao
Liu, Kan
Guan, Haibing
Jiang, Li
author_facet Yang, Ning
Liu, Fangxin
Wang, Junjie
Yang, Tao
Liu, Kan
Guan, Haibing
Jiang, Li
contents Large language models (LLMs) have achieved remarkable performance across a wide range of NLP tasks. However, their substantial inference cost poses a major barrier to real-world deployment, especially in latency-sensitive scenarios. To address this challenge, we propose \textbf{DASH}, an adaptive layer-skipping framework that dynamically selects computation paths conditioned on input characteristics. We model the skipping process as a Markov Decision Process (MDP), enabling fine-grained token-level decisions based on intermediate representations. To mitigate potential performance degradation caused by skipping, we introduce a lightweight compensation mechanism that injects differential rewards into the decision process. Furthermore, we design an asynchronous execution strategy that overlaps layer computation with policy evaluation to minimize runtime overhead. Experiments on multiple LLM architectures and NLP benchmarks show that our method achieves significant inference acceleration while maintaining competitive task performance, outperforming existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17420
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
Yang, Ning
Liu, Fangxin
Wang, Junjie
Yang, Tao
Liu, Kan
Guan, Haibing
Jiang, Li
Computation and Language
Machine Learning
Large language models (LLMs) have achieved remarkable performance across a wide range of NLP tasks. However, their substantial inference cost poses a major barrier to real-world deployment, especially in latency-sensitive scenarios. To address this challenge, we propose \textbf{DASH}, an adaptive layer-skipping framework that dynamically selects computation paths conditioned on input characteristics. We model the skipping process as a Markov Decision Process (MDP), enabling fine-grained token-level decisions based on intermediate representations. To mitigate potential performance degradation caused by skipping, we introduce a lightweight compensation mechanism that injects differential rewards into the decision process. Furthermore, we design an asynchronous execution strategy that overlaps layer computation with policy evaluation to minimize runtime overhead. Experiments on multiple LLM architectures and NLP benchmarks show that our method achieves significant inference acceleration while maintaining competitive task performance, outperforming existing methods.
title DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.17420