Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xuan, Yunyi, Chen, Weijie, Yang, Shicai, Xie, Di, Lin, Luojun, Zhuang, Yueting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
Unveiling Encoder-Free Vision-Language Models
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
MAO: Efficient Model-Agnostic Optimization of Prompt Tuning for Vision-Language Models
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
von: Gao, Ziyuan, et al.
Veröffentlicht: (2025)
von: Gao, Ziyuan, et al.
Veröffentlicht: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
Mitigating Image Captioning Hallucinations in Vision-Language Models
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
ComAlign: Compositional Alignment in Vision-Language Models
von: Abdollah, Ali, et al.
Veröffentlicht: (2024)
von: Abdollah, Ali, et al.
Veröffentlicht: (2024)
PG-Attack: A Precision-Guided Adversarial Attack Framework Against Vision Foundation Models for Autonomous Driving
von: Fu, Jiyuan, et al.
Veröffentlicht: (2024)
von: Fu, Jiyuan, et al.
Veröffentlicht: (2024)
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
von: Zhu, Hongyi, et al.
Veröffentlicht: (2024)
von: Zhu, Hongyi, et al.
Veröffentlicht: (2024)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2026)
Text-Only Data Synthesis for Vision Language Model Training
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
DuoTeach: Dual Role Self-Teaching for Coarse-to-Fine Decision Coordination in Vision--Language Models
von: Yang, Wei, et al.
Veröffentlicht: (2025)
von: Yang, Wei, et al.
Veröffentlicht: (2025)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
von: Tang, Hao, et al.
Veröffentlicht: (2026)
von: Tang, Hao, et al.
Veröffentlicht: (2026)
Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
FreeEnhance: Tuning-Free Image Enhancement via Content-Consistent Noising-and-Denoising Process
von: Luo, Yang, et al.
Veröffentlicht: (2024)
von: Luo, Yang, et al.
Veröffentlicht: (2024)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
von: Lin, Xiang, et al.
Veröffentlicht: (2025)
von: Lin, Xiang, et al.
Veröffentlicht: (2025)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
von: Shen, Kai, et al.
Veröffentlicht: (2024)
von: Shen, Kai, et al.
Veröffentlicht: (2024)
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
von: He, Xiao, et al.
Veröffentlicht: (2025)
von: He, Xiao, et al.
Veröffentlicht: (2025)
Towards Realistic Low-Light Image Enhancement via ISP Driven Data Modeling
von: Wang, Zhihua, et al.
Veröffentlicht: (2025)
von: Wang, Zhihua, et al.
Veröffentlicht: (2025)
Vision-Language Models Learn Super Images for Efficient Partially Relevant Video Retrieval
von: Nishimura, Taichi, et al.
Veröffentlicht: (2023)
von: Nishimura, Taichi, et al.
Veröffentlicht: (2023)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
PathAsst: A Generative Foundation AI Assistant Towards Artificial General Intelligence of Pathology
von: Sun, Yuxuan, et al.
Veröffentlicht: (2023)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2023)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
von: Song, Jiale, et al.
Veröffentlicht: (2026)
von: Song, Jiale, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
von: Xie, Jingjing, et al.
Veröffentlicht: (2024) -
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
von: Dai, Guangyu, et al.
Veröffentlicht: (2025) -
Unveiling Encoder-Free Vision-Language Models
von: Diao, Haiwen, et al.
Veröffentlicht: (2024) -
DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models
von: Li, Haoyang, et al.
Veröffentlicht: (2025) -
MAO: Efficient Model-Agnostic Optimization of Prompt Tuning for Vision-Language Models
von: Li, Haoyang, et al.
Veröffentlicht: (2025)