Multi-matrix Factorization Attention
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916565648146432 |
|---|---|
| author | Hu, Jingcheng Li, Houyi Zhang, Yinmin Wang, Zili Zhou, Shuigeng Zhang, Xiangyu Shum, Heung-Yeung Jiang, Daxin |
| author_facet | Hu, Jingcheng Li, Houyi Zhang, Yinmin Wang, Zili Zhou, Shuigeng Zhang, Xiangyu Shum, Heung-Yeung Jiang, Daxin |
| contents | We propose novel attention architectures, Multi-matrix Factorization Attention (MFA) and MFA-Key-Reuse (MFA-KR). Existing variants for standard Multi-Head Attention (MHA), including SOTA methods like MLA, fail to maintain as strong performance under stringent Key-Value cache (KV cache) constraints. MFA enhances model capacity by efficiently scaling up both the number and dimension of attention heads through low-rank matrix factorization in the Query-Key (QK) circuit. Extending MFA, MFA-KR further reduces memory requirements by repurposing the key cache as value through value projection re-parameterization. MFA's design enables strong model capacity when working under tight KV cache budget, while MFA-KR is suitable for even harsher KV cache limits with minor performance trade-off. Notably, in our extensive and large-scale experiments, the proposed architecture outperforms MLA and performs comparably to MHA, while reducing KV cache usage by up to 56% and 93.7%, respectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_19255 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Multi-matrix Factorization Attention Hu, Jingcheng Li, Houyi Zhang, Yinmin Wang, Zili Zhou, Shuigeng Zhang, Xiangyu Shum, Heung-Yeung Jiang, Daxin Machine Learning Computation and Language We propose novel attention architectures, Multi-matrix Factorization Attention (MFA) and MFA-Key-Reuse (MFA-KR). Existing variants for standard Multi-Head Attention (MHA), including SOTA methods like MLA, fail to maintain as strong performance under stringent Key-Value cache (KV cache) constraints. MFA enhances model capacity by efficiently scaling up both the number and dimension of attention heads through low-rank matrix factorization in the Query-Key (QK) circuit. Extending MFA, MFA-KR further reduces memory requirements by repurposing the key cache as value through value projection re-parameterization. MFA's design enables strong model capacity when working under tight KV cache budget, while MFA-KR is suitable for even harsher KV cache limits with minor performance trade-off. Notably, in our extensive and large-scale experiments, the proposed architecture outperforms MLA and performs comparably to MHA, while reducing KV cache usage by up to 56% and 93.7%, respectively. |
| title | Multi-matrix Factorization Attention |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2412.19255 |