Saved in:
Bibliographic Details
Main Authors: Ma, Xindian, Lu, Yidi, Zhang, Peng, Zhang, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.02197
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912868579934208
author Ma, Xindian
Lu, Yidi
Zhang, Peng
Zhang, Jing
author_facet Ma, Xindian
Lu, Yidi
Zhang, Peng
Zhang, Jing
contents The integration of visual information into Large Language Models (LLMs) has enabled Multimodal LLMs (MLLMs), but the quadratic memory and computational costs of Transformer architectures remain a bottleneck. Existing KV cache eviction strategies fail to address the heterogeneous attention distributions between visual and text tokens, leading to suboptimal efficiency or degraded performance. In this paper, we propose Hierarchical Adaptive Eviction (HAE), a KV cache eviction framework that optimizes text-visual token interaction in MLLMs by implementing Dual-Attention Pruning during pre-filling (leveraging visual token sparsity and attention variance) and a Dynamic Decoding Eviction Strategy (inspired by OS Recycle Bins) during decoding. HAE minimizes KV cache usage across layers, reduces computational overhead via index broadcasting, and theoretically ensures superior information integrity and lower error bounds compared to greedy strategies, enhancing efficiency in both comprehension and generation tasks. Empirically, HAE reduces KV-Cache memory by 41\% with minimal accuracy loss (0.3\% drop) in image understanding tasks and accelerates story generation inference by 1.5x while maintaining output quality on Phi3.5-Vision-Instruct model.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02197
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
Ma, Xindian
Lu, Yidi
Zhang, Peng
Zhang, Jing
Machine Learning
Artificial Intelligence
The integration of visual information into Large Language Models (LLMs) has enabled Multimodal LLMs (MLLMs), but the quadratic memory and computational costs of Transformer architectures remain a bottleneck. Existing KV cache eviction strategies fail to address the heterogeneous attention distributions between visual and text tokens, leading to suboptimal efficiency or degraded performance. In this paper, we propose Hierarchical Adaptive Eviction (HAE), a KV cache eviction framework that optimizes text-visual token interaction in MLLMs by implementing Dual-Attention Pruning during pre-filling (leveraging visual token sparsity and attention variance) and a Dynamic Decoding Eviction Strategy (inspired by OS Recycle Bins) during decoding. HAE minimizes KV cache usage across layers, reduces computational overhead via index broadcasting, and theoretically ensures superior information integrity and lower error bounds compared to greedy strategies, enhancing efficiency in both comprehension and generation tasks. Empirically, HAE reduces KV-Cache memory by 41\% with minimal accuracy loss (0.3\% drop) in image understanding tasks and accelerates story generation inference by 1.5x while maintaining output quality on Phi3.5-Vision-Instruct model.
title Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.02197