FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Kunxi, Xiong, Yufan, Jiang, Zhonghua, Zhou, Yiyun, Wang, Zhaode, Lv, Chengfei, Zhang, Shengyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909900283576320
author Li, Kunxi
Xiong, Yufan
Jiang, Zhonghua
Zhou, Yiyun
Wang, Zhaode
Lv, Chengfei
Zhang, Shengyu
author_facet Li, Kunxi
Xiong, Yufan
Jiang, Zhonghua
Zhou, Yiyun
Wang, Zhaode
Lv, Chengfei
Zhang, Shengyu
contents Traditional KV cache eviction strategies, which discard less critical KV-pairs based on attention scores, often degrade generation quality, causing context loss or hallucinations. Recent efforts shift toward KV merging, merging eviction tokens with retention tokens based on similarity. However, in multimodal scenarios, distributional biases across modality tokens and attentional biases in cross-modal interactions limit its effectiveness. This work introduces FlowMM, an adaptive framework for cross-modal information flow-guided multimodal KV cache merging. FlowMM leverages cross-modal information flow to dynamically apply layer-specific merging strategies, capturing modality-specific patterns while preserving contextual integrity. Furthermore, we introduce a sensitivity-adaptive token matching mechanism that jointly evaluates token similarity and task-critical sensitivity, merging low-risk tokens while safeguarding high-sensitivity ones. Extensive experiments across diverse leading MLLMs show that FlowMM reduces KV cache memory by 80% to 95% and decoding latency by 1.3-1.8x, while maintaining competitive task performance.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05534
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
Li, Kunxi
Xiong, Yufan
Jiang, Zhonghua
Zhou, Yiyun
Wang, Zhaode
Lv, Chengfei
Zhang, Shengyu
Computation and Language
Traditional KV cache eviction strategies, which discard less critical KV-pairs based on attention scores, often degrade generation quality, causing context loss or hallucinations. Recent efforts shift toward KV merging, merging eviction tokens with retention tokens based on similarity. However, in multimodal scenarios, distributional biases across modality tokens and attentional biases in cross-modal interactions limit its effectiveness. This work introduces FlowMM, an adaptive framework for cross-modal information flow-guided multimodal KV cache merging. FlowMM leverages cross-modal information flow to dynamically apply layer-specific merging strategies, capturing modality-specific patterns while preserving contextual integrity. Furthermore, we introduce a sensitivity-adaptive token matching mechanism that jointly evaluates token similarity and task-critical sensitivity, merging low-risk tokens while safeguarding high-sensitivity ones. Extensive experiments across diverse leading MLLMs show that FlowMM reduces KV cache memory by 80% to 95% and decoding latency by 1.3-1.8x, while maintaining competitive task performance.
title FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
topic Computation and Language
url https://arxiv.org/abs/2511.05534