Saved in:
Bibliographic Details
Main Authors: Glorioso, Paolo, Anthony, Quentin, Tokpanov, Yury, Golubeva, Anna, Shyam, Vasudev, Whittington, James, Pilault, Jonathan, Millidge, Beren
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2411.15242
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910710049538048
author Glorioso, Paolo
Anthony, Quentin
Tokpanov, Yury
Golubeva, Anna
Shyam, Vasudev
Whittington, James
Pilault, Jonathan
Millidge, Beren
author_facet Glorioso, Paolo
Anthony, Quentin
Tokpanov, Yury
Golubeva, Anna
Shyam, Vasudev
Whittington, James
Pilault, Jonathan
Millidge, Beren
contents In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance against the leading open-weights models of their class, while achieving substantial gains in inference latency, throughput, and memory efficiency. The Zamba2 series builds upon our initial work with Zamba1-7B, optimizing its architecture, training and annealing datasets, and training for up to three trillion tokens. We provide open-source weights for all models of the Zamba2 series as well as instruction-tuned variants that are strongly competitive against comparable instruct-tuned models of their class. We additionally open-source the pretraining dataset, which we call Zyda-2, used to train the Zamba2 series of models. The models and datasets used in this work are openly available at https://huggingface.co/Zyphra
format Preprint
id arxiv_https___arxiv_org_abs_2411_15242
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Zamba2 Suite: Technical Report
Glorioso, Paolo
Anthony, Quentin
Tokpanov, Yury
Golubeva, Anna
Shyam, Vasudev
Whittington, James
Pilault, Jonathan
Millidge, Beren
Machine Learning
Artificial Intelligence
Computation and Language
In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance against the leading open-weights models of their class, while achieving substantial gains in inference latency, throughput, and memory efficiency. The Zamba2 series builds upon our initial work with Zamba1-7B, optimizing its architecture, training and annealing datasets, and training for up to three trillion tokens. We provide open-source weights for all models of the Zamba2 series as well as instruction-tuned variants that are strongly competitive against comparable instruct-tuned models of their class. We additionally open-source the pretraining dataset, which we call Zyda-2, used to train the Zamba2 series of models. The models and datasets used in this work are openly available at https://huggingface.co/Zyphra
title The Zamba2 Suite: Technical Report
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.15242