Empirical Recipes for Efficient and Compact Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Jiabo, Li, Zhizhong, Sajadmanesh, Sina, Zhuang, Weiming, Lyu, Lingjuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
by: Kang, Weitai, et al.
Published: (2025)
by: Kang, Weitai, et al.
Published: (2025)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
by: Chen, Siyi, et al.
Published: (2026)
by: Chen, Siyi, et al.
Published: (2026)
FedWon: Triumphing Multi-domain Federated Learning Without Normalization
by: Zhuang, Weiming, et al.
Published: (2023)
by: Zhuang, Weiming, et al.
Published: (2023)
A Simple Background Augmentation Method for Object Detection with Diffusion Model
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
by: Huang, Jiabo, et al.
Published: (2025)
by: Huang, Jiabo, et al.
Published: (2025)
FEDEXCHANGE: Bridging the Domain Gap in Federated Object Detection for Free
by: Yuan, Haolin, et al.
Published: (2025)
by: Yuan, Haolin, et al.
Published: (2025)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
by: Yan, Yuping, et al.
Published: (2025)
by: Yan, Yuping, et al.
Published: (2025)
UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
by: Wang, Yimu, et al.
Published: (2025)
by: Wang, Yimu, et al.
Published: (2025)
Towards Fundamentally Scalable Model Selection: Asymptotically Fast Update and Selection
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
CopyJudge: Automated Copyright Infringement Identification and Mitigation in Text-to-Image Diffusion Models
by: Liu, Shunchang, et al.
Published: (2025)
by: Liu, Shunchang, et al.
Published: (2025)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
by: Li, Hongxing, et al.
Published: (2025)
by: Li, Hongxing, et al.
Published: (2025)
ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification
by: He, Yefei, et al.
Published: (2024)
by: He, Yefei, et al.
Published: (2024)
KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
by: Xu, Wanshun, et al.
Published: (2025)
by: Xu, Wanshun, et al.
Published: (2025)
PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models
by: Meng, Yu, et al.
Published: (2025)
by: Meng, Yu, et al.
Published: (2025)
Generalizable Object Re-Identification via Visual In-Context Prompting
by: Huang, Zhizhong, et al.
Published: (2025)
by: Huang, Zhizhong, et al.
Published: (2025)
FoPru: Focal Pruning for Efficient Large Vision-Language Models
by: Jiang, Lei, et al.
Published: (2024)
by: Jiang, Lei, et al.
Published: (2024)
COALA: A Practical and Vision-Centric Federated Learning Platform
by: Zhuang, Weiming, et al.
Published: (2024)
by: Zhuang, Weiming, et al.
Published: (2024)
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
by: Li, Dingming, et al.
Published: (2025)
by: Li, Dingming, et al.
Published: (2025)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
by: Motamed, Saman, et al.
Published: (2025)
by: Motamed, Saman, et al.
Published: (2025)
SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse Autoencoders
by: Li, Qing, et al.
Published: (2025)
by: Li, Qing, et al.
Published: (2025)
Vero: An Open RL Recipe for General Visual Reasoning
by: Sarch, Gabriel, et al.
Published: (2026)
by: Sarch, Gabriel, et al.
Published: (2026)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
by: Huang, Mingzhe, et al.
Published: (2026)
by: Huang, Mingzhe, et al.
Published: (2026)
Is Synthetic Image Useful for Transfer Learning? An Investigation into Data Generation, Volume, and Utilization
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
Are Large Vision Language Models Good Game Players?
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
Landsat30-AU: A Vision-Language Dataset for Australian Landsat Imagery
by: Ma, Sai, et al.
Published: (2025)
by: Ma, Sai, et al.
Published: (2025)
Training-Free Layout-to-Image Generation with Marginal Attention Constraints
by: Chen, Huancheng, et al.
Published: (2024)
by: Chen, Huancheng, et al.
Published: (2024)
Finding needles in a haystack: A Black-Box Approach to Invisible Watermark Detection
by: Pan, Minzhou, et al.
Published: (2024)
by: Pan, Minzhou, et al.
Published: (2024)
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
by: Huang, Yingbing, et al.
Published: (2026)
by: Huang, Yingbing, et al.
Published: (2026)
Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning
by: Busaranuvong, Palawat, et al.
Published: (2026)
by: Busaranuvong, Palawat, et al.
Published: (2026)
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
by: Liu, Xiwei, et al.
Published: (2026)
by: Liu, Xiwei, et al.
Published: (2026)
Probing Perceptual Constancy in Large Vision-Language Models
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
by: Freitas, Miguel Monte e, et al.
Published: (2026)
by: Freitas, Miguel Monte e, et al.
Published: (2026)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
by: Zhao, Xiaobei, et al.
Published: (2025)
by: Zhao, Xiaobei, et al.
Published: (2025)
Efficient Learning for Product Attributes with Compact Multimodal Models
by: Kulkarni, Mandar
Published: (2025)
by: Kulkarni, Mandar
Published: (2025)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
by: Ye, Wei, et al.
Published: (2024)
by: Ye, Wei, et al.
Published: (2024)
Learning to Retrieve Navigable Candidates for Efficient Vision-and-Language Navigation
by: Gu, Shutian, et al.
Published: (2026)
by: Gu, Shutian, et al.
Published: (2026)
Activity Recognition on Avatar-Anonymized Datasets with Masked Differential Privacy
by: Schneider, David, et al.
Published: (2024)
by: Schneider, David, et al.
Published: (2024)
Similar Items
-
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
by: Kang, Weitai, et al.
Published: (2025) -
On the Limits of Token Reduction for Efficient Unified Vision Language Training
by: Chen, Siyi, et al.
Published: (2026) -
FedWon: Triumphing Multi-domain Federated Learning Without Normalization
by: Zhuang, Weiming, et al.
Published: (2023) -
A Simple Background Augmentation Method for Object Detection with Diffusion Model
by: Li, Yuhang, et al.
Published: (2024) -
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)