Hymba: A Hybrid-head Architecture for Small Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Xin, Fu, Yonggan, Diao, Shizhe, Byeon, Wonmin, Chen, Zijia, Mahabaleshwarkar, Ameya Sunil, Liu, Shih-Yang, Van Keirsbilck, Matthijs, Chen, Min-Hung, Suhara, Yoshi, Lin, Yingyan, Kautz, Jan, Molchanov, Pavlo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915028878229504
author Dong, Xin
Fu, Yonggan
Diao, Shizhe
Byeon, Wonmin
Chen, Zijia
Mahabaleshwarkar, Ameya Sunil
Liu, Shih-Yang
Van Keirsbilck, Matthijs
Chen, Min-Hung
Suhara, Yoshi
Lin, Yingyan
Kautz, Jan
Molchanov, Pavlo
author_facet Dong, Xin
Fu, Yonggan
Diao, Shizhe
Byeon, Wonmin
Chen, Zijia
Mahabaleshwarkar, Ameya Sunil
Liu, Shih-Yang
Van Keirsbilck, Matthijs
Chen, Min-Hung
Suhara, Yoshi
Lin, Yingyan
Kautz, Jan
Molchanov, Pavlo
contents We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates transformer attention mechanisms with state space models (SSMs) for enhanced efficiency. Attention heads provide high-resolution recall, while SSM heads enable efficient context summarization. Additionally, we introduce learnable meta tokens that are prepended to prompts, storing critical information and alleviating the "forced-to-attend" burden associated with attention mechanisms. This model is further optimized by incorporating cross-layer key-value (KV) sharing and partial sliding window attention, resulting in a compact cache size. During development, we conducted a controlled study comparing various architectures under identical settings and observed significant advantages of our proposed architecture. Notably, Hymba achieves state-of-the-art results for small LMs: Our Hymba-1.5B-Base model surpasses all sub-2B public models in performance and even outperforms Llama-3.2-3B with 1.32% higher average accuracy, an 11.67x cache size reduction, and 3.49x throughput.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13676
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hymba: A Hybrid-head Architecture for Small Language Models
Dong, Xin
Fu, Yonggan
Diao, Shizhe
Byeon, Wonmin
Chen, Zijia
Mahabaleshwarkar, Ameya Sunil
Liu, Shih-Yang
Van Keirsbilck, Matthijs
Chen, Min-Hung
Suhara, Yoshi
Lin, Yingyan
Kautz, Jan
Molchanov, Pavlo
Computation and Language
Artificial Intelligence
Machine Learning
We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates transformer attention mechanisms with state space models (SSMs) for enhanced efficiency. Attention heads provide high-resolution recall, while SSM heads enable efficient context summarization. Additionally, we introduce learnable meta tokens that are prepended to prompts, storing critical information and alleviating the "forced-to-attend" burden associated with attention mechanisms. This model is further optimized by incorporating cross-layer key-value (KV) sharing and partial sliding window attention, resulting in a compact cache size. During development, we conducted a controlled study comparing various architectures under identical settings and observed significant advantages of our proposed architecture. Notably, Hymba achieves state-of-the-art results for small LMs: Our Hymba-1.5B-Base model surpasses all sub-2B public models in performance and even outperforms Llama-3.2-3B with 1.32% higher average accuracy, an 11.67x cache size reduction, and 3.49x throughput.
title Hymba: A Hybrid-head Architecture for Small Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.13676