[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Ao, Sun, Fengyuan, Chen, Hui, Lin, Zijia, Han, Jungong, Ding, Guiguang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RepViT-SAM: Towards Real-Time Segmenting Anything
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
LSNet: See Large, Focus Small
by: Wang, Ao, et al.
Published: (2025)
by: Wang, Ao, et al.
Published: (2025)
RepViT: Revisiting Mobile CNN From ViT Perspective
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
YOLOE: Real-Time Seeing Anything
by: Wang, Ao, et al.
Published: (2025)
by: Wang, Ao, et al.
Published: (2025)
PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
by: Sun, Fengyuan, et al.
Published: (2025)
by: Sun, Fengyuan, et al.
Published: (2025)
YOLOv10: Real-Time End-to-End Object Detection
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
by: Xiong, Yizhe, et al.
Published: (2025)
by: Xiong, Yizhe, et al.
Published: (2025)
PYRA: Parallel Yielding Re-Activation for Training-Inference Efficient Task Adaptation
by: Xiong, Yizhe, et al.
Published: (2024)
by: Xiong, Yizhe, et al.
Published: (2024)
YOLO-UniOW: Efficient Universal Open-World Object Detection
by: Liu, Lihao, et al.
Published: (2024)
by: Liu, Lihao, et al.
Published: (2024)
Context Enhancement with Reconstruction as Sequence for Unified Unsupervised Anomaly Detection
by: Yang, Hui-Yue, et al.
Published: (2024)
by: Yang, Hui-Yue, et al.
Published: (2024)
Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning
by: Yang, Hui-Yue, et al.
Published: (2024)
by: Yang, Hui-Yue, et al.
Published: (2024)
Learn from the Learnt: Source-Free Active Domain Adaptation via Contrastive Sampling and Visual Persistence
by: Lyu, Mengyao, et al.
Published: (2024)
by: Lyu, Mengyao, et al.
Published: (2024)
PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
by: Sun, Fengyuan, et al.
Published: (2025)
by: Sun, Fengyuan, et al.
Published: (2025)
Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
by: Liang, Yiwen, et al.
Published: (2025)
by: Liang, Yiwen, et al.
Published: (2025)
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
by: Hao, Tianxiang, et al.
Published: (2023)
by: Hao, Tianxiang, et al.
Published: (2023)
Tracking and Segmenting Anything in Any Modality
by: Zhang, Tianlu, et al.
Published: (2025)
by: Zhang, Tianlu, et al.
Published: (2025)
LLMI3D: MLLM-based 3D Perception from a Single 2D Image
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
CASP: Few-Shot Class-Incremental Learning with CLS Token Attention Steering Prompts
by: Huang, Shuai, et al.
Published: (2026)
by: Huang, Shuai, et al.
Published: (2026)
DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
by: Shen, Leqi, et al.
Published: (2025)
by: Shen, Leqi, et al.
Published: (2025)
Quantized Prompt for Efficient Generalization of Vision-Language Models
by: Hao, Tianxiang, et al.
Published: (2024)
by: Hao, Tianxiang, et al.
Published: (2024)
Image Tokenizer Needs Post-Training
by: Qiu, Kai, et al.
Published: (2025)
by: Qiu, Kai, et al.
Published: (2025)
Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification
by: Nguyen, Y Hop, et al.
Published: (2025)
by: Nguyen, Y Hop, et al.
Published: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
by: Duan, Yuxiang, et al.
Published: (2025)
by: Duan, Yuxiang, et al.
Published: (2025)
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
by: Marouani, Alexis, et al.
Published: (2026)
by: Marouani, Alexis, et al.
Published: (2026)
Unlocking [CLS] Features for Continual Post-Training
by: Yildirim, Murat Onur, et al.
Published: (2025)
by: Yildirim, Murat Onur, et al.
Published: (2025)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
by: Shen, Leqi, et al.
Published: (2024)
by: Shen, Leqi, et al.
Published: (2024)
Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning
by: Lyu, Mengyao, et al.
Published: (2025)
by: Lyu, Mengyao, et al.
Published: (2025)
One-Dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing Applications
by: Lyu, Mengyao, et al.
Published: (2023)
by: Lyu, Mengyao, et al.
Published: (2023)
TiCLS : Tightly Coupled Language Text Spotter
by: Jang, Leeje, et al.
Published: (2026)
by: Jang, Leeje, et al.
Published: (2026)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
Grounding Everything in Tokens for Multimodal Large Language Models
by: Ren, Xiangxuan, et al.
Published: (2025)
by: Ren, Xiangxuan, et al.
Published: (2025)
SAM-Body4D: Training-Free 4D Human Body Mesh Recovery from Videos
by: Gao, Mingqi, et al.
Published: (2025)
by: Gao, Mingqi, et al.
Published: (2025)
WaveFace: Authentic Face Restoration with Efficient Frequency Recovery
by: Miao, Yunqi, et al.
Published: (2024)
by: Miao, Yunqi, et al.
Published: (2024)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025)
by: Kim, Sanghwan, et al.
Published: (2025)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
by: Tang, Zihan, et al.
Published: (2026)
by: Tang, Zihan, et al.
Published: (2026)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Similar Items
-
RepViT-SAM: Towards Real-Time Segmenting Anything
by: Wang, Ao, et al.
Published: (2023) -
LSNet: See Large, Focus Small
by: Wang, Ao, et al.
Published: (2025) -
RepViT: Revisiting Mobile CNN From ViT Perspective
by: Wang, Ao, et al.
Published: (2023) -
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
by: Wang, Ao, et al.
Published: (2023) -
YOLOE: Real-Time Seeing Anything
by: Wang, Ao, et al.
Published: (2025)