CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Zhongzhu, Bie, Fengxiang, Chen, Ziyan, Zhang, Zhenyu, Yang, Yibo, Wang, Junxiong, Athiwaratkun, Ben, Wu, Xiaoxia, Song, Shuaiwen Leon
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915872798408704
author Zhou, Zhongzhu
Bie, Fengxiang
Chen, Ziyan
Zhang, Zhenyu
Yang, Yibo
Wang, Junxiong
Athiwaratkun, Ben
Wu, Xiaoxia
Song, Shuaiwen Leon
author_facet Zhou, Zhongzhu
Bie, Fengxiang
Chen, Ziyan
Zhang, Zhenyu
Yang, Yibo
Wang, Junxiong
Athiwaratkun, Ben
Wu, Xiaoxia
Song, Shuaiwen Leon
contents Converting pretrained attention modules such as grouped-query attention (GQA) into multi-head latent attention (MLA) can improve expressivity without increasing KV-cache cost, making it attractive for efficient inference. However, many practical conversion baselines rely on weight-only low-rank approximations (e.g., SVD-style initializations) and uniform rank allocation. They focus on minimizing the difference between weight matrices rather than on how those weights affect input activations, ignore the covariance structure of activations, and enforce uniform rank across layers, causing activation drift and degraded attention fidelity. To address these issues, we propose CARE, a Covariance-Aware, Rank-Enhanced MLA conversion pipeline under a fixed KV width. CARE introduces three key steps: (i) activation-preserving factorization, which aligns the approximation with the actual input activations rather than just the weights; (ii) adjusted-rank allocation, which spreads a fixed KV budget across layers by giving more capacity to layers that need it most; and (iii) KV-parity mapping, which reparameterizes the converted K and V to fit the MLA format while keeping the KV-cache size unchanged. Our method outperforms a uniform-rank SVD baseline on Qwen3-4B/30B-A3B-Instruct-2507 and Llama-3.1-8B/70B-Instruct, reducing one-shot perplexity by up to 215x and improving mean accuracy by up to 1.70x at matched KV budgets. With a brief post-SVD healing fine-tune, we fully recover the original model's accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17946
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
Zhou, Zhongzhu
Bie, Fengxiang
Chen, Ziyan
Zhang, Zhenyu
Yang, Yibo
Wang, Junxiong
Athiwaratkun, Ben
Wu, Xiaoxia
Song, Shuaiwen Leon
Machine Learning
Artificial Intelligence
Converting pretrained attention modules such as grouped-query attention (GQA) into multi-head latent attention (MLA) can improve expressivity without increasing KV-cache cost, making it attractive for efficient inference. However, many practical conversion baselines rely on weight-only low-rank approximations (e.g., SVD-style initializations) and uniform rank allocation. They focus on minimizing the difference between weight matrices rather than on how those weights affect input activations, ignore the covariance structure of activations, and enforce uniform rank across layers, causing activation drift and degraded attention fidelity. To address these issues, we propose CARE, a Covariance-Aware, Rank-Enhanced MLA conversion pipeline under a fixed KV width. CARE introduces three key steps: (i) activation-preserving factorization, which aligns the approximation with the actual input activations rather than just the weights; (ii) adjusted-rank allocation, which spreads a fixed KV budget across layers by giving more capacity to layers that need it most; and (iii) KV-parity mapping, which reparameterizes the converted K and V to fit the MLA format while keeping the KV-cache size unchanged. Our method outperforms a uniform-rank SVD baseline on Qwen3-4B/30B-A3B-Instruct-2507 and Llama-3.1-8B/70B-Instruct, reducing one-shot perplexity by up to 215x and improving mean accuracy by up to 1.70x at matched KV budgets. With a brief post-SVD healing fine-tune, we fully recover the original model's accuracy.
title CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.17946