LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sun, Boyuan, Zhao, Jiaxing, Wei, Xihan, Hou, Qibin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918072522113024
author Sun, Boyuan
Zhao, Jiaxing
Wei, Xihan
Hou, Qibin
author_facet Sun, Boyuan
Zhao, Jiaxing
Wei, Xihan
Hou, Qibin
contents In this paper, we present LLaVA-Scissor, a training-free token compression strategy designed for video multimodal large language models. Previous methods mostly attempt to compress tokens based on attention scores, but fail to effectively capture all semantic regions and often lead to token redundancy. Differently, we propose to leverage the Semantic Connected Components (SCC) approach that assigns tokens to distinct semantic regions within the token set, ensuring comprehensive semantic coverage. The outcome is a two-step spatio-temporal token compression strategy that utilizes SCC in both spatial and temporal domains. This strategy can effectively compress tokens by representing the entire video with a set of non-overlapping semantic tokens. We conduct extensive evaluations of the token compression capabilities of LLaVA-Scissor across diverse video understanding benchmarks, including video question answering, long video understanding, and comprehensive multi-choices benchmarks. Experimental results show that the proposed LLaVA-Scissor outperforms other token compression methods, achieving superior performance in various video understanding benchmarks, particularly at low token retention ratios. Project page: https://github.com/HumanMLLM/LLaVA-Scissor.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
Sun, Boyuan
Zhao, Jiaxing
Wei, Xihan
Hou, Qibin
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Multimedia
In this paper, we present LLaVA-Scissor, a training-free token compression strategy designed for video multimodal large language models. Previous methods mostly attempt to compress tokens based on attention scores, but fail to effectively capture all semantic regions and often lead to token redundancy. Differently, we propose to leverage the Semantic Connected Components (SCC) approach that assigns tokens to distinct semantic regions within the token set, ensuring comprehensive semantic coverage. The outcome is a two-step spatio-temporal token compression strategy that utilizes SCC in both spatial and temporal domains. This strategy can effectively compress tokens by representing the entire video with a set of non-overlapping semantic tokens. We conduct extensive evaluations of the token compression capabilities of LLaVA-Scissor across diverse video understanding benchmarks, including video question answering, long video understanding, and comprehensive multi-choices benchmarks. Experimental results show that the proposed LLaVA-Scissor outperforms other token compression methods, achieving superior performance in various video understanding benchmarks, particularly at low token retention ratios. Project page: https://github.com/HumanMLLM/LLaVA-Scissor.
title LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Multimedia
url https://arxiv.org/abs/2506.21862