Pipelined Decoder for Efficient Context-Aware Text Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Zixian, Niu, Chenxu, Gu, Yu, Xiao, Gengyang, Huang, Xinwei, Cheng, Gong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909667884531712
author Huang, Zixian
Niu, Chenxu
Gu, Yu
Xiao, Gengyang
Huang, Xinwei
Cheng, Gong
author_facet Huang, Zixian
Niu, Chenxu
Gu, Yu
Xiao, Gengyang
Huang, Xinwei
Cheng, Gong
contents As the basis of generative AI, an autoregressive model requires the generation of a new token depending on all the previously generated tokens, which brings high quality but also restricts the model to generate tokens one by one, forming a bottleneck limiting the generation speed. In this paper, we propose a new decoder architecture that efficiently generates text in parallel for context-aware generation tasks. Our proposed pipelined decoder initiates the generation of multiple subsequences simultaneously, and, at each time-step, it generates a new token for each subsequence to realize parallelism. Experiments on multiple text generation tasks, including question answering, text summarization, and keyphrase generation, show that our pipelined decoder significantly improves the generation speed without a significant loss of generation quality or additional memory consumption.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23431
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pipelined Decoder for Efficient Context-Aware Text Generation
Huang, Zixian
Niu, Chenxu
Gu, Yu
Xiao, Gengyang
Huang, Xinwei
Cheng, Gong
Computation and Language
Artificial Intelligence
As the basis of generative AI, an autoregressive model requires the generation of a new token depending on all the previously generated tokens, which brings high quality but also restricts the model to generate tokens one by one, forming a bottleneck limiting the generation speed. In this paper, we propose a new decoder architecture that efficiently generates text in parallel for context-aware generation tasks. Our proposed pipelined decoder initiates the generation of multiple subsequences simultaneously, and, at each time-step, it generates a new token for each subsequence to realize parallelism. Experiments on multiple text generation tasks, including question answering, text summarization, and keyphrase generation, show that our pipelined decoder significantly improves the generation speed without a significant loss of generation quality or additional memory consumption.
title Pipelined Decoder for Efficient Context-Aware Text Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.23431