Dodo: Dynamic Contextual Compression for Decoder-only LMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Qin, Guanghui, Rosset, Corby, Chau, Ethan C., Rao, Nikhil, Van Durme, Benjamin
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913600288849920
author Qin, Guanghui
Rosset, Corby
Chau, Ethan C.
Rao, Nikhil
Van Durme, Benjamin
author_facet Qin, Guanghui
Rosset, Corby
Chau, Ethan C.
Rao, Nikhil
Van Durme, Benjamin
contents Transformer-based language models (LMs) are inefficient in long contexts. We propose Dodo, a solution for context compression. Instead of one vector per token in a standard transformer model, Dodo represents text with a dynamic number of hidden states at each layer, reducing the cost of self-attention to a fraction of typical time and space. Moreover, off-the-shelf models such as LLaMA can be adapted to Dodo by efficient parameter tuning methods such as LoRA. In use, Dodo can act as either an autoregressive LM or a context compressor for downstream tasks. We demonstrate through experiments in language modeling, question answering, and summarization that Dodo retains capabilities in these tasks, while drastically reducing the overhead during decoding. For example, in the autoencoding task, Dodo shrinks context at a 20x compression ratio with a BLEU score of 98% for reconstruction, achieving nearly lossless encoding.
format Preprint
id arxiv_https___arxiv_org_abs_2310_02409
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Dodo: Dynamic Contextual Compression for Decoder-only LMs
Qin, Guanghui
Rosset, Corby
Chau, Ethan C.
Rao, Nikhil
Van Durme, Benjamin
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; I.2.6
Transformer-based language models (LMs) are inefficient in long contexts. We propose Dodo, a solution for context compression. Instead of one vector per token in a standard transformer model, Dodo represents text with a dynamic number of hidden states at each layer, reducing the cost of self-attention to a fraction of typical time and space. Moreover, off-the-shelf models such as LLaMA can be adapted to Dodo by efficient parameter tuning methods such as LoRA. In use, Dodo can act as either an autoregressive LM or a context compressor for downstream tasks. We demonstrate through experiments in language modeling, question answering, and summarization that Dodo retains capabilities in these tasks, while drastically reducing the overhead during decoding. For example, in the autoencoding task, Dodo shrinks context at a 20x compression ratio with a BLEU score of 98% for reconstruction, achieving nearly lossless encoding.
title Dodo: Dynamic Contextual Compression for Decoder-only LMs
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; I.2.6
url https://arxiv.org/abs/2310.02409