RecycleGPT: An Autoregressive Language Model with Recyclable Module

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Yufan, He, Qiaozhi, Zhuang, Xiaomin, Wu, Zhihua, Wang, Kunpeng, Zhao, Wenlai, Yang, Guangwen
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909208422645760
author Jiang, Yufan
He, Qiaozhi
Zhuang, Xiaomin
Wu, Zhihua
Wang, Kunpeng
Zhao, Wenlai
Yang, Guangwen
author_facet Jiang, Yufan
He, Qiaozhi
Zhuang, Xiaomin
Wu, Zhihua
Wang, Kunpeng
Zhao, Wenlai
Yang, Guangwen
contents Existing large language models have to run K times to generate a sequence of K tokens. In this paper, we present RecycleGPT, a generative language model with fast decoding speed by recycling pre-generated model states without running the whole model in multiple steps. Our approach relies on the observation that adjacent tokens in a sequence usually have strong correlations and the next token in a sequence can be reasonably guessed or inferred based on the preceding ones. Experiments and analysis demonstrate the effectiveness of our approach in lowering inference latency, achieving up to 1.4x speedup while preserving high performance.
format Preprint
id arxiv_https___arxiv_org_abs_2308_03421
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle RecycleGPT: An Autoregressive Language Model with Recyclable Module
Jiang, Yufan
He, Qiaozhi
Zhuang, Xiaomin
Wu, Zhihua
Wang, Kunpeng
Zhao, Wenlai
Yang, Guangwen
Computation and Language
Artificial Intelligence
Existing large language models have to run K times to generate a sequence of K tokens. In this paper, we present RecycleGPT, a generative language model with fast decoding speed by recycling pre-generated model states without running the whole model in multiple steps. Our approach relies on the observation that adjacent tokens in a sequence usually have strong correlations and the next token in a sequence can be reasonably guessed or inferred based on the preceding ones. Experiments and analysis demonstrate the effectiveness of our approach in lowering inference latency, achieving up to 1.4x speedup while preserving high performance.
title RecycleGPT: An Autoregressive Language Model with Recyclable Module
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2308.03421