Saved in:
Bibliographic Details
Main Authors: Luo, Xuan, Wang, Weizhi, Yan, Xifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.23798
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914082019344384
author Luo, Xuan
Wang, Weizhi
Yan, Xifeng
author_facet Luo, Xuan
Wang, Weizhi
Yan, Xifeng
contents Various layer-skipping methods have been proposed to accelerate token generation in large language models (LLMs). However, limited attention has been paid to a fundamental question: How do computational demands vary across the generation of different tokens? In this work, we introduce FlexiDepth, a method that dynamically adjusts the number of Transformer layers used in text generation. By incorporating a plug-in router and adapter, FlexiDepth enables adaptive computation in LLMs without modifying their original parameters. Applied to Llama-3-8B, it skips 8 out of 32 layers while maintaining full benchmark performance. Our experiments reveal that computational demands in LLMs significantly vary based on token type. Specifically, generating repetitive tokens or fixed phrases requires fewer layers, whereas producing tokens involving computation or high uncertainty requires more layers. Despite the computational savings, FlexiDepth does not yet achieve wall-clock speedup due to varied skipping patterns and I/O overhead. To inspire future work and advance research on practical speedup, we open-sourced FlexiDepth and a dataset documenting its layer allocation patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23798
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Layer-skipping in Pre-trained LLMs
Luo, Xuan
Wang, Weizhi
Yan, Xifeng
Computation and Language
Artificial Intelligence
Various layer-skipping methods have been proposed to accelerate token generation in large language models (LLMs). However, limited attention has been paid to a fundamental question: How do computational demands vary across the generation of different tokens? In this work, we introduce FlexiDepth, a method that dynamically adjusts the number of Transformer layers used in text generation. By incorporating a plug-in router and adapter, FlexiDepth enables adaptive computation in LLMs without modifying their original parameters. Applied to Llama-3-8B, it skips 8 out of 32 layers while maintaining full benchmark performance. Our experiments reveal that computational demands in LLMs significantly vary based on token type. Specifically, generating repetitive tokens or fixed phrases requires fewer layers, whereas producing tokens involving computation or high uncertainty requires more layers. Despite the computational savings, FlexiDepth does not yet achieve wall-clock speedup due to varied skipping patterns and I/O overhead. To inspire future work and advance research on practical speedup, we open-sourced FlexiDepth and a dataset documenting its layer allocation patterns.
title Adaptive Layer-skipping in Pre-trained LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.23798