MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Sixun, Hu, Juhua, Zhang, Mian, Yin, Ming, Fu, Yanjie, Qian, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Model Efficiency: Multi-Agent Inference with Large Models
by: Dong, Sixun, et al.
Published: (2026)
by: Dong, Sixun, et al.
Published: (2026)
Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
by: Dong, Sixun, et al.
Published: (2025)
by: Dong, Sixun, et al.
Published: (2025)
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
Online Zero-Shot Classification with CLIP
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering
by: Yao, Jiawei, et al.
Published: (2024)
by: Yao, Jiawei, et al.
Published: (2024)
SeA: Semantic Adversarial Augmentation for Last Layer Features from Unsupervised Representation Learning
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
Text-Guided Mixup Towards Long-Tailed Image Categorization
by: Franklin, Richard, et al.
Published: (2024)
by: Franklin, Richard, et al.
Published: (2024)
Dual-disentangled Deep Multiple Clustering
by: Yao, Jiawei, et al.
Published: (2024)
by: Yao, Jiawei, et al.
Published: (2024)
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
by: Jadhav, Avadhoot, et al.
Published: (2025)
by: Jadhav, Avadhoot, et al.
Published: (2025)
CLGRPO: Reasoning Ability Enhancement for Small VLMs
by: Wang, Fanyi, et al.
Published: (2025)
by: Wang, Fanyi, et al.
Published: (2025)
MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
by: Chen, Qinyu, et al.
Published: (2025)
by: Chen, Qinyu, et al.
Published: (2025)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
by: Lee, Jaehoon, et al.
Published: (2026)
by: Lee, Jaehoon, et al.
Published: (2026)
T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition
by: Yeh, Chen, et al.
Published: (2024)
by: Yeh, Chen, et al.
Published: (2024)
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
by: Zhang, Xike, et al.
Published: (2026)
by: Zhang, Xike, et al.
Published: (2026)
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
by: Aich, Abhishek, et al.
Published: (2026)
by: Aich, Abhishek, et al.
Published: (2026)
Clapper: Compact Learning and Video Representation in VLMs
by: Kong, Lingyu, et al.
Published: (2025)
by: Kong, Lingyu, et al.
Published: (2025)
Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
by: Xie, Zhuofan, et al.
Published: (2026)
by: Xie, Zhuofan, et al.
Published: (2026)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
by: Cho, Hyeonwoo, et al.
Published: (2026)
by: Cho, Hyeonwoo, et al.
Published: (2026)
Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation
by: Joo, Minseok, et al.
Published: (2026)
by: Joo, Minseok, et al.
Published: (2026)
GeoDecoder: Empowering Multimodal Map Understanding
by: Qi, Feng, et al.
Published: (2024)
by: Qi, Feng, et al.
Published: (2024)
ApET: Approximation-Error Guided Token Compression for Efficient VLMs
by: Ma, Qiankun, et al.
Published: (2026)
by: Ma, Qiankun, et al.
Published: (2026)
CityLLaVA: Efficient Fine-Tuning for VLMs in City Scenario
by: Duan, Zhizhao, et al.
Published: (2024)
by: Duan, Zhizhao, et al.
Published: (2024)
GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
by: Zhang, Wenchuan, et al.
Published: (2025)
by: Zhang, Wenchuan, et al.
Published: (2025)
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
by: Zhong, Qihuang, et al.
Published: (2026)
by: Zhong, Qihuang, et al.
Published: (2026)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
by: Ye, Maoyuan, et al.
Published: (2025)
by: Ye, Maoyuan, et al.
Published: (2025)
PolarMAE: Efficient Fetal Ultrasound Pre-training via Semantic Screening and Polar-Guided Masking
by: Lv, Meng, et al.
Published: (2026)
by: Lv, Meng, et al.
Published: (2026)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)
by: Zhang, Xintong, et al.
Published: (2025)
Deep Pre-Alignment for VLMs
by: Yu, Tianyu, et al.
Published: (2026)
by: Yu, Tianyu, et al.
Published: (2026)
Data Factory with Minimal Human Effort Using VLMs
by: Ye, Jiaojiao, et al.
Published: (2025)
by: Ye, Jiaojiao, et al.
Published: (2025)
Rethinking the Need for Source Models: Source-Free Domain Adaptation from Scratch Guided by a Vision-Language Model
by: Bingtao, Zhou, et al.
Published: (2026)
by: Bingtao, Zhou, et al.
Published: (2026)
One-shot Optimized Steering Vector for Hallucination Mitigation for VLMs
by: Shi, Youxu, et al.
Published: (2026)
by: Shi, Youxu, et al.
Published: (2026)
MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs
by: Ma, Junpeng, et al.
Published: (2025)
by: Ma, Junpeng, et al.
Published: (2025)
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
by: Liu, Ting, et al.
Published: (2024)
by: Liu, Ting, et al.
Published: (2024)
VLMs Guided Interpretable Decision Making for Autonomous Driving
by: Hu, Xin, et al.
Published: (2025)
by: Hu, Xin, et al.
Published: (2025)
QwenSafe: Multimodal Content Rating Description Identification via Preference-Aligned VLMs
by: Denipitiyage, Dishanika, et al.
Published: (2026)
by: Denipitiyage, Dishanika, et al.
Published: (2026)
Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks
by: Cai, Jinjin, et al.
Published: (2024)
by: Cai, Jinjin, et al.
Published: (2024)
Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers
by: Wang, Jiancheng, et al.
Published: (2026)
by: Wang, Jiancheng, et al.
Published: (2026)
Similar Items
-
Rethinking Model Efficiency: Multi-Agent Inference with Large Models
by: Dong, Sixun, et al.
Published: (2026) -
Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
by: Dong, Sixun, et al.
Published: (2025) -
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
by: Qian, Qi, et al.
Published: (2024) -
Online Zero-Shot Classification with CLIP
by: Qian, Qi, et al.
Published: (2024) -
Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering
by: Yao, Jiawei, et al.
Published: (2024)