Hymba: A Hybrid-head Architecture for Small Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915028878229504 |
|---|---|
| author | Dong, Xin Fu, Yonggan Diao, Shizhe Byeon, Wonmin Chen, Zijia Mahabaleshwarkar, Ameya Sunil Liu, Shih-Yang Van Keirsbilck, Matthijs Chen, Min-Hung Suhara, Yoshi Lin, Yingyan Kautz, Jan Molchanov, Pavlo |
| author_facet | Dong, Xin Fu, Yonggan Diao, Shizhe Byeon, Wonmin Chen, Zijia Mahabaleshwarkar, Ameya Sunil Liu, Shih-Yang Van Keirsbilck, Matthijs Chen, Min-Hung Suhara, Yoshi Lin, Yingyan Kautz, Jan Molchanov, Pavlo |
| contents | We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates transformer attention mechanisms with state space models (SSMs) for enhanced efficiency. Attention heads provide high-resolution recall, while SSM heads enable efficient context summarization. Additionally, we introduce learnable meta tokens that are prepended to prompts, storing critical information and alleviating the "forced-to-attend" burden associated with attention mechanisms. This model is further optimized by incorporating cross-layer key-value (KV) sharing and partial sliding window attention, resulting in a compact cache size. During development, we conducted a controlled study comparing various architectures under identical settings and observed significant advantages of our proposed architecture. Notably, Hymba achieves state-of-the-art results for small LMs: Our Hymba-1.5B-Base model surpasses all sub-2B public models in performance and even outperforms Llama-3.2-3B with 1.32% higher average accuracy, an 11.67x cache size reduction, and 3.49x throughput. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_13676 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Hymba: A Hybrid-head Architecture for Small Language Models Dong, Xin Fu, Yonggan Diao, Shizhe Byeon, Wonmin Chen, Zijia Mahabaleshwarkar, Ameya Sunil Liu, Shih-Yang Van Keirsbilck, Matthijs Chen, Min-Hung Suhara, Yoshi Lin, Yingyan Kautz, Jan Molchanov, Pavlo Computation and Language Artificial Intelligence Machine Learning We propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates transformer attention mechanisms with state space models (SSMs) for enhanced efficiency. Attention heads provide high-resolution recall, while SSM heads enable efficient context summarization. Additionally, we introduce learnable meta tokens that are prepended to prompts, storing critical information and alleviating the "forced-to-attend" burden associated with attention mechanisms. This model is further optimized by incorporating cross-layer key-value (KV) sharing and partial sliding window attention, resulting in a compact cache size. During development, we conducted a controlled study comparing various architectures under identical settings and observed significant advantages of our proposed architecture. Notably, Hymba achieves state-of-the-art results for small LMs: Our Hymba-1.5B-Base model surpasses all sub-2B public models in performance and even outperforms Llama-3.2-3B with 1.32% higher average accuracy, an 11.67x cache size reduction, and 3.49x throughput. |
| title | Hymba: A Hybrid-head Architecture for Small Language Models |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2411.13676 |