COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Jinqi, Yin, Miao, Gong, Yu, Zang, Xiao, Ren, Jian, Yuan, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ELRT: Efficient Low-Rank Training for Compact Convolutional Neural Networks
by: Sui, Yang, et al.
Published: (2024)
by: Sui, Yang, et al.
Published: (2024)
HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
by: Xiao, Jinqi, et al.
Published: (2023)
by: Xiao, Jinqi, et al.
Published: (2023)
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
by: Yang, Cheng, et al.
Published: (2026)
by: Yang, Cheng, et al.
Published: (2026)
TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models
by: Zhong, Jing, et al.
Published: (2025)
by: Zhong, Jing, et al.
Published: (2025)
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
by: Yu, Binhe, et al.
Published: (2025)
by: Yu, Binhe, et al.
Published: (2025)
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
by: Apedo, Yvon, et al.
Published: (2026)
by: Apedo, Yvon, et al.
Published: (2026)
Spiking Vision Transformer with Saccadic Attention
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
DEFormer: DCT-driven Enhancement Transformer for Low-light Image and Dark Vision
by: Yin, Xiangchen, et al.
Published: (2023)
by: Yin, Xiangchen, et al.
Published: (2023)
Vision Augmentation Prediction Autoencoder with Attention Design (VAPAAD)
by: Yin, Yiqiao
Published: (2024)
by: Yin, Yiqiao
Published: (2024)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models
by: Arif, Kazi Hasan Ibn, et al.
Published: (2024)
by: Arif, Kazi Hasan Ibn, et al.
Published: (2024)
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
Do Pre-trained Vision-Language Models Encode Object States?
by: Newman, Kaleb, et al.
Published: (2024)
by: Newman, Kaleb, et al.
Published: (2024)
Lite-Mind: Towards Efficient and Robust Brain Representation Network
by: Gong, Zixuan, et al.
Published: (2023)
by: Gong, Zixuan, et al.
Published: (2023)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
by: Zou, Dongyun, et al.
Published: (2026)
by: Zou, Dongyun, et al.
Published: (2026)
Towards Lossless Ultimate Vision Token Compression for VLMs
by: Zheng, Dehua, et al.
Published: (2025)
by: Zheng, Dehua, et al.
Published: (2025)
COAP: Memory-Efficient Training with Correlation-Aware Gradient Projection
by: Xiao, Jinqi, et al.
Published: (2024)
by: Xiao, Jinqi, et al.
Published: (2024)
DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models
by: Wang, JiYang, et al.
Published: (2026)
by: Wang, JiYang, et al.
Published: (2026)
Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering
by: Li, Zhifei, et al.
Published: (2025)
by: Li, Zhifei, et al.
Published: (2025)
Vision-Language Models in Remote Sensing: Current Progress and Future Trends
by: Li, Xiang, et al.
Published: (2023)
by: Li, Xiang, et al.
Published: (2023)
Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
by: Deniz, Omer Faruk, et al.
Published: (2026)
by: Deniz, Omer Faruk, et al.
Published: (2026)
LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model
by: Xue, Yuan, et al.
Published: (2024)
by: Xue, Yuan, et al.
Published: (2024)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
by: Zou, Siyu, et al.
Published: (2024)
by: Zou, Siyu, et al.
Published: (2024)
Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
by: Zhang, Naifu, et al.
Published: (2025)
by: Zhang, Naifu, et al.
Published: (2025)
Prompt-Agnostic Adversarial Perturbation for Customized Diffusion Models
by: Wan, Cong, et al.
Published: (2024)
by: Wan, Cong, et al.
Published: (2024)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
by: Huang, Yingbing, et al.
Published: (2026)
by: Huang, Yingbing, et al.
Published: (2026)
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
by: Zhang, Shu-Hao, et al.
Published: (2025)
by: Zhang, Shu-Hao, et al.
Published: (2025)
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
by: Wang, Huanyu, et al.
Published: (2025)
by: Wang, Huanyu, et al.
Published: (2025)
Efficient Online Continual Learning in Sensor-Based Human Activity Recognition
by: Zhang, Yao, et al.
Published: (2025)
by: Zhang, Yao, et al.
Published: (2025)
Online Handwritten Signature Verification Based on Temporal-Spatial Graph Attention Transformer
by: Yuan, Hai-jie, et al.
Published: (2025)
by: Yuan, Hai-jie, et al.
Published: (2025)
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
by: Böhle, Moritz, et al.
Published: (2025)
by: Böhle, Moritz, et al.
Published: (2025)
A-VL: Adaptive Attention for Large Vision-Language Models
by: Zhang, Junyang, et al.
Published: (2024)
by: Zhang, Junyang, et al.
Published: (2024)
InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation
by: Xiao, Jinqi, et al.
Published: (2025)
by: Xiao, Jinqi, et al.
Published: (2025)
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
by: Qian, Chen, et al.
Published: (2026)
by: Qian, Chen, et al.
Published: (2026)
Customize Segment Anything Model for Multi-Modal Semantic Segmentation with Mixture of LoRA Experts
by: Zhu, Chenyang, et al.
Published: (2024)
by: Zhu, Chenyang, et al.
Published: (2024)
Task-Aware KV Compression For Cost-Effective Long Video Understanding
by: Qin, Minghao, et al.
Published: (2025)
by: Qin, Minghao, et al.
Published: (2025)
Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
by: Zhang, Enming, et al.
Published: (2025)
by: Zhang, Enming, et al.
Published: (2025)
AdaRing: Towards Ultra-Light Vision-Language Adaptation via Cross-Layer Tensor Ring Decomposition
by: Huang, Ying, et al.
Published: (2025)
by: Huang, Ying, et al.
Published: (2025)
Similar Items
-
ELRT: Efficient Low-Rank Training for Compact Convolutional Neural Networks
by: Sui, Yang, et al.
Published: (2024) -
HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
by: Xiao, Jinqi, et al.
Published: (2023) -
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
by: Yang, Cheng, et al.
Published: (2026) -
TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model
by: Yang, Cheng, et al.
Published: (2025) -
ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models
by: Zhong, Jing, et al.
Published: (2025)