VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zaiwei, Meyer, Gregory P., Lu, Zhichao, Shrivastava, Ashish, Ravichandran, Avinash, Wolff, Eric M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
VLMine: Long-Tail Data Mining with Vision Language Models
by: Ye, Mao, et al.
Published: (2024)
by: Ye, Mao, et al.
Published: (2024)
Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality
by: Chen, Liyan, et al.
Published: (2024)
by: Chen, Liyan, et al.
Published: (2024)
InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models
by: Saini, Nirat, et al.
Published: (2024)
by: Saini, Nirat, et al.
Published: (2024)
Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation
by: Bafghi, Reza Akbarian, et al.
Published: (2025)
by: Bafghi, Reza Akbarian, et al.
Published: (2025)
BD-KD: Balancing the Divergences for Online Knowledge Distillation
by: Amara, Ibtihel, et al.
Published: (2022)
by: Amara, Ibtihel, et al.
Published: (2022)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
Feature Fusion from Head to Tail for Long-Tailed Visual Recognition
by: Li, Mengke, et al.
Published: (2023)
by: Li, Mengke, et al.
Published: (2023)
VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization
by: Waheed, Sania, et al.
Published: (2025)
by: Waheed, Sania, et al.
Published: (2025)
Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models
by: Sun, Haoyi, et al.
Published: (2026)
by: Sun, Haoyi, et al.
Published: (2026)
MST-KD: Multiple Specialized Teachers Knowledge Distillation for Fair Face Recognition
by: Caldeira, Eduarda, et al.
Published: (2024)
by: Caldeira, Eduarda, et al.
Published: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
by: Xue, Xizhe, et al.
Published: (2024)
by: Xue, Xizhe, et al.
Published: (2024)
Long-Tailed Visual Recognition via Permutation-Invariant Head-to-Tail Feature Fusion
by: Li, Mengke, et al.
Published: (2025)
by: Li, Mengke, et al.
Published: (2025)
Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
by: Chen, Tao, et al.
Published: (2025)
by: Chen, Tao, et al.
Published: (2025)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
by: Huang, Zhipeng, et al.
Published: (2024)
by: Huang, Zhipeng, et al.
Published: (2024)
XR-VLM: Cross-Relationship Modeling with Multi-part Prompts and Visual Features for Fine-Grained Recognition
by: Wang, Chuanming, et al.
Published: (2025)
by: Wang, Chuanming, et al.
Published: (2025)
SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
by: Khaki, Samir, et al.
Published: (2025)
by: Khaki, Samir, et al.
Published: (2025)
Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual Recognition
by: Li, Mengke, et al.
Published: (2024)
by: Li, Mengke, et al.
Published: (2024)
TopKD: Top-scaled Knowledge Distillation
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Long-Tailed Continual Learning For Visual Food Recognition
by: He, Jiangpeng, et al.
Published: (2023)
by: He, Jiangpeng, et al.
Published: (2023)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)
by: Fateh, Fawad Javed, et al.
Published: (2024)
FreeKD: Knowledge Distillation via Semantic Frequency Prompt
by: Zhang, Yuan, et al.
Published: (2023)
by: Zhang, Yuan, et al.
Published: (2023)
CogVLM: Visual Expert for Pretrained Language Models
by: Wang, Weihan, et al.
Published: (2023)
by: Wang, Weihan, et al.
Published: (2023)
TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
by: Stergiou, Alexandros
Published: (2025)
by: Stergiou, Alexandros
Published: (2025)
GenMM: Geometrically and Temporally Consistent Multimodal Data Generation for Video and LiDAR
by: Singh, Bharat, et al.
Published: (2024)
by: Singh, Bharat, et al.
Published: (2024)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points Sampling
by: Wang, Yu, et al.
Published: (2022)
by: Wang, Yu, et al.
Published: (2022)
SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge
by: He, Yumeng, et al.
Published: (2025)
by: He, Yumeng, et al.
Published: (2025)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026)
by: Zhang, Wanyue, et al.
Published: (2026)
MoKD: Multi-Task Optimization for Knowledge Distillation
by: Hayder, Zeeshan, et al.
Published: (2025)
by: Hayder, Zeeshan, et al.
Published: (2025)
EA-KD: Entropy-based Adaptive Knowledge Distillation
by: Su, Chi-Ping, et al.
Published: (2023)
by: Su, Chi-Ping, et al.
Published: (2023)
PersonaVLM: Long-Term Personalized Multimodal LLMs
by: Nie, Chang, et al.
Published: (2026)
by: Nie, Chang, et al.
Published: (2026)
DocVLM: Make Your VLM an Efficient Reader
by: Nacson, Mor Shpigel, et al.
Published: (2024)
by: Nacson, Mor Shpigel, et al.
Published: (2024)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
by: Yu, Seonghoon, et al.
Published: (2026)
by: Yu, Seonghoon, et al.
Published: (2026)
IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models
by: Shi, Liang, et al.
Published: (2026)
by: Shi, Liang, et al.
Published: (2026)
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
by: Qiu, Jason, et al.
Published: (2026)
by: Qiu, Jason, et al.
Published: (2026)
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration
by: Han, Jiayi, et al.
Published: (2025)
by: Han, Jiayi, et al.
Published: (2025)
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
by: Zhang, Zijian, et al.
Published: (2025)
by: Zhang, Zijian, et al.
Published: (2025)
Similar Items
-
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
by: Xu, Yi, et al.
Published: (2024) -
VLMine: Long-Tail Data Mining with Vision Language Models
by: Ye, Mao, et al.
Published: (2024) -
Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality
by: Chen, Liyan, et al.
Published: (2024) -
InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models
by: Saini, Nirat, et al.
Published: (2024) -
Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation
by: Bafghi, Reza Akbarian, et al.
Published: (2025)