Zamba: A Compact 7B SSM Hybrid Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Glorioso, Paolo, Anthony, Quentin, Tokpanov, Yury, Whittington, James, Pilault, Jonathan, Ibrahim, Adam, Millidge, Beren
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913364177846272
author Glorioso, Paolo
Anthony, Quentin
Tokpanov, Yury
Whittington, James
Pilault, Jonathan
Ibrahim, Adam
Millidge, Beren
author_facet Glorioso, Paolo
Anthony, Quentin
Tokpanov, Yury
Whittington, James
Pilault, Jonathan
Ibrahim, Adam
Millidge, Beren
contents In this technical report, we present Zamba, a novel 7B SSM-transformer hybrid model which achieves competitive performance against leading open-weight models at a comparable scale. Zamba is trained on 1T tokens from openly available datasets and is the best non-transformer model at this scale. Zamba pioneers a unique architecture combining a Mamba backbone with a single shared attention module, thus obtaining the benefits of attention at minimal parameter cost. Due to its architecture, Zamba is significantly faster at inference than comparable transformer models and requires substantially less memory for generation of long sequences. Zamba is pretrained in two phases: the first phase is based on existing web datasets, while the second one consists of annealing the model over high-quality instruct and synthetic datasets, and is characterized by a rapid learning rate decay. We open-source the weights and all checkpoints for Zamba, through both phase 1 and annealing phases.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16712
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zamba: A Compact 7B SSM Hybrid Model
Glorioso, Paolo
Anthony, Quentin
Tokpanov, Yury
Whittington, James
Pilault, Jonathan
Ibrahim, Adam
Millidge, Beren
Machine Learning
Artificial Intelligence
Computation and Language
In this technical report, we present Zamba, a novel 7B SSM-transformer hybrid model which achieves competitive performance against leading open-weight models at a comparable scale. Zamba is trained on 1T tokens from openly available datasets and is the best non-transformer model at this scale. Zamba pioneers a unique architecture combining a Mamba backbone with a single shared attention module, thus obtaining the benefits of attention at minimal parameter cost. Due to its architecture, Zamba is significantly faster at inference than comparable transformer models and requires substantially less memory for generation of long sequences. Zamba is pretrained in two phases: the first phase is based on existing web datasets, while the second one consists of annealing the model over high-quality instruct and synthetic datasets, and is characterized by a rapid learning rate decay. We open-source the weights and all checkpoints for Zamba, through both phase 1 and annealing phases.
title Zamba: A Compact 7B SSM Hybrid Model
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.16712