LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Hongyu, Zhang, Feng, Jin, Wenwei, Hu, Huanling, Shi, Tianjun, Jiang, Shikai, Hu, Yao, Li, Jiawei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911687753334784
author Lu, Hongyu
Zhang, Feng
Jin, Wenwei
Hu, Huanling
Shi, Tianjun
Jiang, Shikai
Hu, Yao
Li, Jiawei
author_facet Lu, Hongyu
Zhang, Feng
Jin, Wenwei
Hu, Huanling
Shi, Tianjun
Jiang, Shikai
Hu, Yao
Li, Jiawei
contents Large vision-language models (LVLMs) achieve strong multimodal understanding, but their inference cost grows rapidly with the number of visual tokens, especially for high-resolution images and long videos. Existing attention-based methods estimate token importance from attention scores, which may introduce positional bias, while representation-based methods reduce visual redundancy based on feature relations or reconstruction errors, overlooking the global structure of the visual token set. In this paper, we revisit visual token compression from the perspective of low-rank compressibility. Across models and datasets, we observe that visual token representations exhibit a pronounced low-rank structure, with a dominant subspace that remains stable even after a large fraction of tokens is randomly removed. Motivated by this finding, we propose LRCP, a training-free compression framework that first estimates the dominant low-rank subspace of visual tokens via PCA, and then scores each token by its projection residual onto this subspace, retaining tokens that are poorly explained by the low-rank background. Extensive experiments show that LRCP achieves superior results, preserving 94.7% of the original image-understanding performance with an 88.9% token reduction and 97.8% of the average video-understanding accuracy with an 87.5% token reduction.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15621
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs
Lu, Hongyu
Zhang, Feng
Jin, Wenwei
Hu, Huanling
Shi, Tianjun
Jiang, Shikai
Hu, Yao
Li, Jiawei
Computer Vision and Pattern Recognition
68T45, 68Q25, 65F55
I.2.10; I.4.8; I.5.4
Large vision-language models (LVLMs) achieve strong multimodal understanding, but their inference cost grows rapidly with the number of visual tokens, especially for high-resolution images and long videos. Existing attention-based methods estimate token importance from attention scores, which may introduce positional bias, while representation-based methods reduce visual redundancy based on feature relations or reconstruction errors, overlooking the global structure of the visual token set. In this paper, we revisit visual token compression from the perspective of low-rank compressibility. Across models and datasets, we observe that visual token representations exhibit a pronounced low-rank structure, with a dominant subspace that remains stable even after a large fraction of tokens is randomly removed. Motivated by this finding, we propose LRCP, a training-free compression framework that first estimates the dominant low-rank subspace of visual tokens via PCA, and then scores each token by its projection residual onto this subspace, retaining tokens that are poorly explained by the low-rank background. Extensive experiments show that LRCP achieves superior results, preserving 94.7% of the original image-understanding performance with an 88.9% token reduction and 97.8% of the average video-understanding accuracy with an 87.5% token reduction.
title LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs
topic Computer Vision and Pattern Recognition
68T45, 68Q25, 65F55
I.2.10; I.4.8; I.5.4
url https://arxiv.org/abs/2605.15621