Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jang, Young Kyun, Kim, Donghyun, Lim, Ser-nam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
MATE: Meet At The Embedding -- Connecting Images with Long Texts
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
UVIS: Unsupervised Video Instance Segmentation
von: Huang, Shuaiyi, et al.
Veröffentlicht: (2024)
von: Huang, Shuaiyi, et al.
Veröffentlicht: (2024)
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
von: Han, Gyuwon, et al.
Veröffentlicht: (2026)
von: Han, Gyuwon, et al.
Veröffentlicht: (2026)
Body-Hand Modality Expertized Networks with Cross-attention for Fine-grained Skeleton Action Recognition
von: Cho, Seungyeon, et al.
Veröffentlicht: (2025)
von: Cho, Seungyeon, et al.
Veröffentlicht: (2025)
BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion
von: Xiang, Sike, et al.
Veröffentlicht: (2025)
von: Xiang, Sike, et al.
Veröffentlicht: (2025)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
von: He, Bo, et al.
Veröffentlicht: (2024)
von: He, Bo, et al.
Veröffentlicht: (2024)
Cross-Modal Adapter for Vision-Language Retrieval
von: Jiang, Haojun, et al.
Veröffentlicht: (2022)
von: Jiang, Haojun, et al.
Veröffentlicht: (2022)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
Adaptive Teaching with Shared Classifier for Knowledge Distillation
von: Jang, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Jang, Jaeyeon, et al.
Veröffentlicht: (2024)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2026)
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2026)
GOAL: Global-local Object Alignment Learning
von: Choi, Hyungyu, et al.
Veröffentlicht: (2025)
von: Choi, Hyungyu, et al.
Veröffentlicht: (2025)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
von: Kim, SiWoo, et al.
Veröffentlicht: (2025)
von: Kim, SiWoo, et al.
Veröffentlicht: (2025)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2025)
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2025)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
von: Feng, Qianhan, et al.
Veröffentlicht: (2024)
von: Feng, Qianhan, et al.
Veröffentlicht: (2024)
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
Generalist Multi-Class Anomaly Detection via Distillation to Two Heterogeneous Student Networks
von: Park, Hangil, et al.
Veröffentlicht: (2025)
von: Park, Hangil, et al.
Veröffentlicht: (2025)
Towards Chunk-Wise Generation for Long Videos
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
von: Bakkali, Souhail, et al.
Veröffentlicht: (2023)
von: Bakkali, Souhail, et al.
Veröffentlicht: (2023)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
Co-learning Single-Step Diffusion Upsampler and Downsampler with Two Discriminators and Distillation
von: Kim, Sohwi, et al.
Veröffentlicht: (2024)
von: Kim, Sohwi, et al.
Veröffentlicht: (2024)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Object-level Self-Distillation for Vision Pretraining
von: Hızlı, Çağlar, et al.
Veröffentlicht: (2025)
von: Hızlı, Çağlar, et al.
Veröffentlicht: (2025)
Cross-Modal Attention Guided Unlearning in Vision-Language Models
von: Bhaila, Karuna, et al.
Veröffentlicht: (2025)
von: Bhaila, Karuna, et al.
Veröffentlicht: (2025)
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data
von: Chung, Dahyun, et al.
Veröffentlicht: (2025)
von: Chung, Dahyun, et al.
Veröffentlicht: (2025)
4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
von: Claessens, Cris, et al.
Veröffentlicht: (2025)
von: Claessens, Cris, et al.
Veröffentlicht: (2025)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
von: Fragomeni, Adriano, et al.
Veröffentlicht: (2025)
von: Fragomeni, Adriano, et al.
Veröffentlicht: (2025)
Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
Enriching Knowledge Distillation with Cross-Modal Teacher Fusion
von: Mansourian, Amir M., et al.
Veröffentlicht: (2025)
von: Mansourian, Amir M., et al.
Veröffentlicht: (2025)
Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency
von: Wei, Riling, et al.
Veröffentlicht: (2025)
von: Wei, Riling, et al.
Veröffentlicht: (2025)
VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models
von: Wang, Chenyu, et al.
Veröffentlicht: (2026)
von: Wang, Chenyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024) -
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024) -
Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024) -
MATE: Meet At The Embedding -- Connecting Images with Long Texts
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024) -
UVIS: Unsupervised Video Instance Segmentation
von: Huang, Shuaiyi, et al.
Veröffentlicht: (2024)