EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Xuchen, Chen, Yanxi, Li, Yaliang, Ding, Bolin, Zhou, Jingren
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917580310052864
author Pan, Xuchen
Chen, Yanxi
Li, Yaliang
Ding, Bolin
Zhou, Jingren
author_facet Pan, Xuchen
Chen, Yanxi
Li, Yaliang
Ding, Bolin
Zhou, Jingren
contents This work introduces EE-Tuning, a lightweight and economical solution to training/tuning early-exit large language models (LLMs). In contrast to the common approach of full-parameter pre-training, EE-Tuning augments any pre-trained (and possibly fine-tuned) standard LLM with additional early-exit layers that are tuned in a parameter-efficient manner, which requires significantly less computational resources and training data. Our implementation of EE-Tuning achieves outstanding training efficiency via extensive performance optimizations, as well as scalability due to its full compatibility with 3D parallelism. Results of systematic experiments validate the efficacy of EE-Tuning, confirming that effective early-exit LLM inference can be achieved with a limited training budget. In hope of making early-exit LLMs accessible to the community, we release the source code of our implementation of EE-Tuning at https://github.com/pan-x-c/EE-LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2402_00518
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models
Pan, Xuchen
Chen, Yanxi
Li, Yaliang
Ding, Bolin
Zhou, Jingren
Machine Learning
Artificial Intelligence
Computation and Language
This work introduces EE-Tuning, a lightweight and economical solution to training/tuning early-exit large language models (LLMs). In contrast to the common approach of full-parameter pre-training, EE-Tuning augments any pre-trained (and possibly fine-tuned) standard LLM with additional early-exit layers that are tuned in a parameter-efficient manner, which requires significantly less computational resources and training data. Our implementation of EE-Tuning achieves outstanding training efficiency via extensive performance optimizations, as well as scalability due to its full compatibility with 3D parallelism. Results of systematic experiments validate the efficacy of EE-Tuning, confirming that effective early-exit LLM inference can be achieved with a limited training budget. In hope of making early-exit LLMs accessible to the community, we release the source code of our implementation of EE-Tuning at https://github.com/pan-x-c/EE-LLM.
title EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2402.00518