Taipan: Efficient and Expressive State Space Language Models with Selective Attention

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Van Nguyen, Chien, Nguyen, Huy Huu, Pham, Thang M., Zhang, Ruiyi, Deilamsalehy, Hanieh, Mathur, Puneet, Rossi, Ryan A., Bui, Trung, Lai, Viet Dac, Dernoncourt, Franck, Nguyen, Thien Huu
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913561821839360
author Van Nguyen, Chien
Nguyen, Huy Huu
Pham, Thang M.
Zhang, Ruiyi
Deilamsalehy, Hanieh
Mathur, Puneet
Rossi, Ryan A.
Bui, Trung
Lai, Viet Dac
Dernoncourt, Franck
Nguyen, Thien Huu
author_facet Van Nguyen, Chien
Nguyen, Huy Huu
Pham, Thang M.
Zhang, Ruiyi
Deilamsalehy, Hanieh
Mathur, Puneet
Rossi, Ryan A.
Bui, Trung
Lai, Viet Dac
Dernoncourt, Franck
Nguyen, Thien Huu
contents Efficient long-context language modeling remains a significant challenge in Natural Language Processing (NLP). While Transformers dominate language tasks, they struggle with long sequences due to quadratic computational complexity in training and linearly scaling memory costs during inference. Recent State Space Models (SSMs) such as Mamba offer alternatives with constant memory usage, but they underperform in tasks requiring extensive in-context retrieval. We introduce Taipan, a novel hybrid architecture that combines Mamba-2 with Selective Attention Layers (SALs). These SALs identify tokens requiring long-range interactions, remove less important features, and then augment their representations using the attention module. This approach balances Mamba's efficiency with Transformer-like performance in memory-intensive tasks. By constraining the attention budget, Taipan extends accurate predictions to context lengths of up to 1 million tokens while preserving computational efficiency. Our experiments demonstrate Taipan's superior performance across various scales and tasks, offering a promising solution for efficient long-context language modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18572
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Taipan: Efficient and Expressive State Space Language Models with Selective Attention
Van Nguyen, Chien
Nguyen, Huy Huu
Pham, Thang M.
Zhang, Ruiyi
Deilamsalehy, Hanieh
Mathur, Puneet
Rossi, Ryan A.
Bui, Trung
Lai, Viet Dac
Dernoncourt, Franck
Nguyen, Thien Huu
Computation and Language
Artificial Intelligence
Machine Learning
Efficient long-context language modeling remains a significant challenge in Natural Language Processing (NLP). While Transformers dominate language tasks, they struggle with long sequences due to quadratic computational complexity in training and linearly scaling memory costs during inference. Recent State Space Models (SSMs) such as Mamba offer alternatives with constant memory usage, but they underperform in tasks requiring extensive in-context retrieval. We introduce Taipan, a novel hybrid architecture that combines Mamba-2 with Selective Attention Layers (SALs). These SALs identify tokens requiring long-range interactions, remove less important features, and then augment their representations using the attention module. This approach balances Mamba's efficiency with Transformer-like performance in memory-intensive tasks. By constraining the attention budget, Taipan extends accurate predictions to context lengths of up to 1 million tokens while preserving computational efficiency. Our experiments demonstrate Taipan's superior performance across various scales and tasks, offering a promising solution for efficient long-context language modeling.
title Taipan: Efficient and Expressive State Space Language Models with Selective Attention
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.18572