CASP: Compression of Large Multimodal Models Based on Attention Sparsity
Fuente:
arXiv
Saved in:
| Main Authors: | Gholami, Mohsen, Akbari, Mohammad, Cannons, Kevin, Zhang, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CPPO: Contrastive Perception Policy Optimization for VLM Agents
by: Rezaei, Ahmad, et al.
Published: (2026)
by: Rezaei, Ahmad, et al.
Published: (2026)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
by: Gholami, Mohsen, et al.
Published: (2025)
by: Gholami, Mohsen, et al.
Published: (2025)
DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
by: Alvar, Saeed Ranjbar, et al.
Published: (2025)
by: Alvar, Saeed Ranjbar, et al.
Published: (2025)
CASP: Few-Shot Class-Incremental Learning with CLS Token Attention Steering Prompts
by: Huang, Shuai, et al.
Published: (2026)
by: Huang, Shuai, et al.
Published: (2026)
Beyond Overall Accuracy: Pose- and Occlusion-driven Fairness Analysis in Pedestrian Detection for Autonomous Driving
by: Khoshkdahan, Mohammad, et al.
Published: (2025)
by: Khoshkdahan, Mohammad, et al.
Published: (2025)
State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models
by: Kim, Geewook, et al.
Published: (2025)
by: Kim, Geewook, et al.
Published: (2025)
LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models
by: Zheng, Pengcheng, et al.
Published: (2026)
by: Zheng, Pengcheng, et al.
Published: (2026)
A Survey of Token Compression for Efficient Multimodal Large Language Models
by: Shao, Kele, et al.
Published: (2025)
by: Shao, Kele, et al.
Published: (2025)
Towards Secure and Usable 3D Assets: A Novel Framework for Automatic Visible Watermarking
by: Singh, Gursimran, et al.
Published: (2024)
by: Singh, Gursimran, et al.
Published: (2024)
Can Multimodal Large Language Models be Guided to Improve Industrial Anomaly Detection?
by: Chen, Zhiling, et al.
Published: (2025)
by: Chen, Zhiling, et al.
Published: (2025)
Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
by: Peng, Tianfan, et al.
Published: (2025)
by: Peng, Tianfan, et al.
Published: (2025)
Y-CA-Net: A Convolutional Attention Based Network for Volumetric Medical Image Segmentation
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
by: Luo, Jiayi, et al.
Published: (2026)
by: Luo, Jiayi, et al.
Published: (2026)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
by: Ye, Hancheng, et al.
Published: (2024)
by: Ye, Hancheng, et al.
Published: (2024)
Understanding and Harnessing Sparsity in Unified Multimodal Models
by: He, Shwai, et al.
Published: (2025)
by: He, Shwai, et al.
Published: (2025)
Sparsity Meets Similarity: Leveraging Long-Tail Distribution for Dynamic Optimized Token Representation in Multimodal Large Language Models
by: Yu, Gaotong, et al.
Published: (2024)
by: Yu, Gaotong, et al.
Published: (2024)
TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
by: Tan, Xudong, et al.
Published: (2025)
by: Tan, Xudong, et al.
Published: (2025)
LaWa: Using Latent Space for In-Generation Image Watermarking
by: Rezaei, Ahmad, et al.
Published: (2024)
by: Rezaei, Ahmad, et al.
Published: (2024)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
by: Fang, Tongcheng, et al.
Published: (2026)
by: Fang, Tongcheng, et al.
Published: (2026)
LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
by: Hu, Qingqiao, et al.
Published: (2025)
by: Hu, Qingqiao, et al.
Published: (2025)
PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models
by: Huang, Mouxiao, et al.
Published: (2025)
by: Huang, Mouxiao, et al.
Published: (2025)
Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention
by: Qiao, Bingtian, et al.
Published: (2026)
by: Qiao, Bingtian, et al.
Published: (2026)
DiTFastAttn: Attention Compression for Diffusion Transformer Models
by: Yuan, Zhihang, et al.
Published: (2024)
by: Yuan, Zhihang, et al.
Published: (2024)
Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
by: Deniz, Omer Faruk, et al.
Published: (2026)
by: Deniz, Omer Faruk, et al.
Published: (2026)
Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completion
by: Xue, Yu, et al.
Published: (2026)
by: Xue, Yu, et al.
Published: (2026)
Linear Attention Modeling for Learned Image Compression
by: Feng, Donghui, et al.
Published: (2025)
by: Feng, Donghui, et al.
Published: (2025)
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
by: Sinha, Sanchit, et al.
Published: (2026)
by: Sinha, Sanchit, et al.
Published: (2026)
VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
by: Xu, Hengbo, et al.
Published: (2026)
by: Xu, Hengbo, et al.
Published: (2026)
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
by: Lei, Lei, et al.
Published: (2025)
by: Lei, Lei, et al.
Published: (2025)
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
by: Tong, Bo, et al.
Published: (2024)
by: Tong, Bo, et al.
Published: (2024)
Learnable Sparsity for Vision Generative Models
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Fast and Controllable Post-training Sparsity: Learning Optimal Sparsity Allocation with Global Constraint in Minutes
by: Gong, Ruihao, et al.
Published: (2024)
by: Gong, Ruihao, et al.
Published: (2024)
Blind Source Separation Based on Sparsity
by: Li, Zhongxuan
Published: (2025)
by: Li, Zhongxuan
Published: (2025)
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
by: Fan, Yingqi, et al.
Published: (2026)
by: Fan, Yingqi, et al.
Published: (2026)
SINR: Sparsity Driven Compressed Implicit Neural Representations
by: Jayasundara, Dhananjaya, et al.
Published: (2025)
by: Jayasundara, Dhananjaya, et al.
Published: (2025)
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
by: Li, Qirui, et al.
Published: (2025)
by: Li, Qirui, et al.
Published: (2025)
Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
by: Bohy, Hugo, et al.
Published: (2025)
by: Bohy, Hugo, et al.
Published: (2025)
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
by: Liu, Juntao, et al.
Published: (2025)
by: Liu, Juntao, et al.
Published: (2025)
LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
by: Hu, Lianyu, et al.
Published: (2025)
by: Hu, Lianyu, et al.
Published: (2025)
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
by: Yao, Linli, et al.
Published: (2024)
by: Yao, Linli, et al.
Published: (2024)
Similar Items
-
CPPO: Contrastive Perception Policy Optimization for VLM Agents
by: Rezaei, Ahmad, et al.
Published: (2026) -
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
by: Gholami, Mohsen, et al.
Published: (2025) -
DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
by: Alvar, Saeed Ranjbar, et al.
Published: (2025) -
CASP: Few-Shot Class-Incremental Learning with CLS Token Attention Steering Prompts
by: Huang, Shuai, et al.
Published: (2026) -
Beyond Overall Accuracy: Pose- and Occlusion-driven Fairness Analysis in Pedestrian Detection for Autonomous Driving
by: Khoshkdahan, Mohammad, et al.
Published: (2025)