Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Zhengyao, Lyu, Pengyuan, Zhang, Chengquan, Lu, Guangming, Yu, Jun, Pei, Wenjie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recognition-Synergistic Scene Text Editing
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025)
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025)
WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-only Supervised Text Spotting
von: Wu, Jingjing, et al.
Veröffentlicht: (2024)
von: Wu, Jingjing, et al.
Veröffentlicht: (2024)
Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity
von: Fang, Zhengyao, et al.
Veröffentlicht: (2026)
von: Fang, Zhengyao, et al.
Veröffentlicht: (2026)
IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs
von: Tan, Yifan, et al.
Veröffentlicht: (2026)
von: Tan, Yifan, et al.
Veröffentlicht: (2026)
Towards Lossless Ultimate Vision Token Compression for VLMs
von: Zheng, Dehua, et al.
Veröffentlicht: (2025)
von: Zheng, Dehua, et al.
Veröffentlicht: (2025)
D2Pruner: Debiased Importance and Structural Diversity for MLLM Token Pruning
von: Zhang, Evelyn, et al.
Veröffentlicht: (2025)
von: Zhang, Evelyn, et al.
Veröffentlicht: (2025)
HIVTP: A Training-Free Method to Improve VLMs Efficiency via Hierarchical Visual Token Pruning Using Middle-Layer-Based Importance Score
von: Xu, Jingqi, et al.
Veröffentlicht: (2025)
von: Xu, Jingqi, et al.
Veröffentlicht: (2025)
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
von: Lee, Yuna, et al.
Veröffentlicht: (2026)
von: Lee, Yuna, et al.
Veröffentlicht: (2026)
Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
von: Ma, Ji, et al.
Veröffentlicht: (2025)
von: Ma, Ji, et al.
Veröffentlicht: (2025)
Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance
von: Liang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Liang, Yuxuan, et al.
Veröffentlicht: (2025)
GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models
von: Pei, Ruiguang, et al.
Veröffentlicht: (2025)
von: Pei, Ruiguang, et al.
Veröffentlicht: (2025)
D$^2$ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition
von: Pei, Wenjie, et al.
Veröffentlicht: (2023)
von: Pei, Wenjie, et al.
Veröffentlicht: (2023)
Video-ToC: Video Tree-of-Cue Reasoning
von: Tan, Qizhong, et al.
Veröffentlicht: (2026)
von: Tan, Qizhong, et al.
Veröffentlicht: (2026)
EditInfinity: Image Editing with Binary-Quantized Generative Models
von: Wang, Jiahuan, et al.
Veröffentlicht: (2025)
von: Wang, Jiahuan, et al.
Veröffentlicht: (2025)
ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
von: Li, Duo, et al.
Veröffentlicht: (2025)
von: Li, Duo, et al.
Veröffentlicht: (2025)
ReDiPrune: Relevance-Diversity Pre-Projection Token Pruning for Efficient Multimodal LLMs
von: Yu, An, et al.
Veröffentlicht: (2026)
von: Yu, An, et al.
Veröffentlicht: (2026)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
von: Bagrov, Natan, et al.
Veröffentlicht: (2025)
von: Bagrov, Natan, et al.
Veröffentlicht: (2025)
HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models
von: Zhu, Qihui, et al.
Veröffentlicht: (2026)
von: Zhu, Qihui, et al.
Veröffentlicht: (2026)
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
von: Mao, Junzhu, et al.
Veröffentlicht: (2025)
von: Mao, Junzhu, et al.
Veröffentlicht: (2025)
Beyond Heuristics: Learnable Density Control for 3D Gaussian Splatting
von: Ning, Zhenhua, et al.
Veröffentlicht: (2026)
von: Ning, Zhenhua, et al.
Veröffentlicht: (2026)
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
von: Liu, Jizhihui, et al.
Veröffentlicht: (2025)
von: Liu, Jizhihui, et al.
Veröffentlicht: (2025)
A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2025)
von: Zeng, Quan-Sheng, et al.
Veröffentlicht: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
ApET: Approximation-Error Guided Token Compression for Efficient VLMs
von: Ma, Qiankun, et al.
Veröffentlicht: (2026)
von: Ma, Qiankun, et al.
Veröffentlicht: (2026)
DiffTrans: Differentiable Geometry-Materials Decomposition for Reconstructing Transparent Objects
von: Li, Changpu, et al.
Veröffentlicht: (2026)
von: Li, Changpu, et al.
Veröffentlicht: (2026)
Attention Debiasing for Token Pruning in Vision Language Models
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
Structural Pruning via Spatial-aware Information Redundancy for Semantic Segmentation
von: Wu, Dongyue, et al.
Veröffentlicht: (2024)
von: Wu, Dongyue, et al.
Veröffentlicht: (2024)
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
Learning Compatible Multi-Prize Subnetworks for Asymmetric Retrieval
von: Sun, Yushuai, et al.
Veröffentlicht: (2025)
von: Sun, Yushuai, et al.
Veröffentlicht: (2025)
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
von: Cao, Jianjian, et al.
Veröffentlicht: (2024)
von: Cao, Jianjian, et al.
Veröffentlicht: (2024)
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
von: Lyu, Pengyuan, et al.
Veröffentlicht: (2024)
von: Lyu, Pengyuan, et al.
Veröffentlicht: (2024)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
Comb, Prune, Distill: Towards Unified Pruning for Vision Model Compression
von: Schmitt, Jonas, et al.
Veröffentlicht: (2024)
von: Schmitt, Jonas, et al.
Veröffentlicht: (2024)
Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
von: Xu, Zhaoqi, et al.
Veröffentlicht: (2025)
von: Xu, Zhaoqi, et al.
Veröffentlicht: (2025)
Diversity-aware Channel Pruning for StyleGAN Compression
von: Chung, Jiwoo, et al.
Veröffentlicht: (2024)
von: Chung, Jiwoo, et al.
Veröffentlicht: (2024)
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Recognition-Synergistic Scene Text Editing
von: Fang, Zhengyao, et al.
Veröffentlicht: (2025) -
WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-only Supervised Text Spotting
von: Wu, Jingjing, et al.
Veröffentlicht: (2024) -
Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity
von: Fang, Zhengyao, et al.
Veröffentlicht: (2026) -
IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs
von: Tan, Yifan, et al.
Veröffentlicht: (2026) -
Towards Lossless Ultimate Vision Token Compression for VLMs
von: Zheng, Dehua, et al.
Veröffentlicht: (2025)