Multi-matrix Factorization Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Jingcheng, Li, Houyi, Zhang, Yinmin, Wang, Zili, Zhou, Shuigeng, Zhang, Xiangyu, Shum, Heung-Yeung, Jiang, Daxin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916565648146432
author Hu, Jingcheng
Li, Houyi
Zhang, Yinmin
Wang, Zili
Zhou, Shuigeng
Zhang, Xiangyu
Shum, Heung-Yeung
Jiang, Daxin
author_facet Hu, Jingcheng
Li, Houyi
Zhang, Yinmin
Wang, Zili
Zhou, Shuigeng
Zhang, Xiangyu
Shum, Heung-Yeung
Jiang, Daxin
contents We propose novel attention architectures, Multi-matrix Factorization Attention (MFA) and MFA-Key-Reuse (MFA-KR). Existing variants for standard Multi-Head Attention (MHA), including SOTA methods like MLA, fail to maintain as strong performance under stringent Key-Value cache (KV cache) constraints. MFA enhances model capacity by efficiently scaling up both the number and dimension of attention heads through low-rank matrix factorization in the Query-Key (QK) circuit. Extending MFA, MFA-KR further reduces memory requirements by repurposing the key cache as value through value projection re-parameterization. MFA's design enables strong model capacity when working under tight KV cache budget, while MFA-KR is suitable for even harsher KV cache limits with minor performance trade-off. Notably, in our extensive and large-scale experiments, the proposed architecture outperforms MLA and performs comparably to MHA, while reducing KV cache usage by up to 56% and 93.7%, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19255
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-matrix Factorization Attention
Hu, Jingcheng
Li, Houyi
Zhang, Yinmin
Wang, Zili
Zhou, Shuigeng
Zhang, Xiangyu
Shum, Heung-Yeung
Jiang, Daxin
Machine Learning
Computation and Language
We propose novel attention architectures, Multi-matrix Factorization Attention (MFA) and MFA-Key-Reuse (MFA-KR). Existing variants for standard Multi-Head Attention (MHA), including SOTA methods like MLA, fail to maintain as strong performance under stringent Key-Value cache (KV cache) constraints. MFA enhances model capacity by efficiently scaling up both the number and dimension of attention heads through low-rank matrix factorization in the Query-Key (QK) circuit. Extending MFA, MFA-KR further reduces memory requirements by repurposing the key cache as value through value projection re-parameterization. MFA's design enables strong model capacity when working under tight KV cache budget, while MFA-KR is suitable for even harsher KV cache limits with minor performance trade-off. Notably, in our extensive and large-scale experiments, the proposed architecture outperforms MLA and performs comparably to MHA, while reducing KV cache usage by up to 56% and 93.7%, respectively.
title Multi-matrix Factorization Attention
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2412.19255