Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.15242 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910710049538048 |
|---|---|
| author | Glorioso, Paolo Anthony, Quentin Tokpanov, Yury Golubeva, Anna Shyam, Vasudev Whittington, James Pilault, Jonathan Millidge, Beren |
| author_facet | Glorioso, Paolo Anthony, Quentin Tokpanov, Yury Golubeva, Anna Shyam, Vasudev Whittington, James Pilault, Jonathan Millidge, Beren |
| contents | In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance against the leading open-weights models of their class, while achieving substantial gains in inference latency, throughput, and memory efficiency. The Zamba2 series builds upon our initial work with Zamba1-7B, optimizing its architecture, training and annealing datasets, and training for up to three trillion tokens. We provide open-source weights for all models of the Zamba2 series as well as instruction-tuned variants that are strongly competitive against comparable instruct-tuned models of their class. We additionally open-source the pretraining dataset, which we call Zyda-2, used to train the Zamba2 series of models. The models and datasets used in this work are openly available at https://huggingface.co/Zyphra |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_15242 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | The Zamba2 Suite: Technical Report Glorioso, Paolo Anthony, Quentin Tokpanov, Yury Golubeva, Anna Shyam, Vasudev Whittington, James Pilault, Jonathan Millidge, Beren Machine Learning Artificial Intelligence Computation and Language In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance against the leading open-weights models of their class, while achieving substantial gains in inference latency, throughput, and memory efficiency. The Zamba2 series builds upon our initial work with Zamba1-7B, optimizing its architecture, training and annealing datasets, and training for up to three trillion tokens. We provide open-source weights for all models of the Zamba2 series as well as instruction-tuned variants that are strongly competitive against comparable instruct-tuned models of their class. We additionally open-source the pretraining dataset, which we call Zyda-2, used to train the Zamba2 series of models. The models and datasets used in this work are openly available at https://huggingface.co/Zyphra |
| title | The Zamba2 Suite: Technical Report |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2411.15242 |