COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Sanghwan, Xiao, Rui, Georgescu, Mariana-Iuliana, Alaniz, Stephan, Akata, Zeynep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLAIR: VLM with Fine-grained Language-informed Image Representations
by: Xiao, Rui, et al.
Published: (2024)
by: Xiao, Rui, et al.
Published: (2024)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025)
by: Kim, Sanghwan, et al.
Published: (2025)
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
by: Xiao, Rui, et al.
Published: (2026)
by: Xiao, Rui, et al.
Published: (2026)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024)
by: Hummel, Thomas, et al.
Published: (2024)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
by: Kim, Jae Myung, et al.
Published: (2025)
by: Kim, Jae Myung, et al.
Published: (2025)
Road Obstacle Video Segmentation
by: Rai, Shyam Nandan, et al.
Published: (2025)
by: Rai, Shyam Nandan, et al.
Published: (2025)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
by: Wu, Boyong, et al.
Published: (2026)
by: Wu, Boyong, et al.
Published: (2026)
X-Aligner: Composed Visual Retrieval without the Bells and Whistles
by: Zheng, Yuqian, et al.
Published: (2026)
by: Zheng, Yuqian, et al.
Published: (2026)
Explaining CLIP Zero-shot Predictions Through Concepts
by: Ozdemir, Onat, et al.
Published: (2026)
by: Ozdemir, Onat, et al.
Published: (2026)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
Anatomical Structure-Guided Medical Vision-Language Pre-training
by: Li, Qingqiu, et al.
Published: (2024)
by: Li, Qingqiu, et al.
Published: (2024)
DataDream: Few-shot Guided Dataset Generation
by: Kim, Jae Myung, et al.
Published: (2024)
by: Kim, Jae Myung, et al.
Published: (2024)
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
by: Huang, Yiran, et al.
Published: (2025)
by: Huang, Yiran, et al.
Published: (2025)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
VLP: A Survey on Vision-Language Pre-training
by: Chen, Feilong, et al.
Published: (2022)
by: Chen, Feilong, et al.
Published: (2022)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
MemLoRA: Distilling Expert Adapters for On-Device Memory Systems
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
Cross-Modal Adapter for Vision-Language Retrieval
by: Jiang, Haojun, et al.
Published: (2022)
by: Jiang, Haojun, et al.
Published: (2022)
Weight Copy and Low-Rank Adaptation for Few-Shot Distillation of Vision Transformers
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
by: Liang, Zhengyang, et al.
Published: (2024)
by: Liang, Zhengyang, et al.
Published: (2024)
Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models
by: Kurzendörfer, David, et al.
Published: (2024)
by: Kurzendörfer, David, et al.
Published: (2024)
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
by: Huang, Qidong, et al.
Published: (2024)
by: Huang, Qidong, et al.
Published: (2024)
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023)
by: Karthik, Shyamgopal, et al.
Published: (2023)
Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images
by: Mahanta, Cristina, et al.
Published: (2025)
by: Mahanta, Cristina, et al.
Published: (2025)
ETHER: Efficient Finetuning of Large-Scale Models with Hyperplane Reflections
by: Bini, Massimo, et al.
Published: (2024)
by: Bini, Massimo, et al.
Published: (2024)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
by: Chen, Junyi, et al.
Published: (2023)
by: Chen, Junyi, et al.
Published: (2023)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
by: Jiang, Lei, et al.
Published: (2025)
by: Jiang, Lei, et al.
Published: (2025)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
by: Zhu, Tinghui, et al.
Published: (2024)
by: Zhu, Tinghui, et al.
Published: (2024)
Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
by: Dalmonte, Francesco, et al.
Published: (2025)
by: Dalmonte, Francesco, et al.
Published: (2025)
LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition
by: Lungu-Stan, Vlad-Constantin, et al.
Published: (2026)
by: Lungu-Stan, Vlad-Constantin, et al.
Published: (2026)
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
by: Gao, Jianjun, et al.
Published: (2024)
by: Gao, Jianjun, et al.
Published: (2024)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
by: Csizmadia, Daniel, et al.
Published: (2025)
by: Csizmadia, Daniel, et al.
Published: (2025)
Superpixel Semantics Representation and Pre-training for Vision-Language Task
by: Zhang, Siyu, et al.
Published: (2023)
by: Zhang, Siyu, et al.
Published: (2023)
Pseudo-Prompt Generating in Pre-trained Vision-Language Models for Multi-Label Medical Image Classification
by: Ye, Yaoqin, et al.
Published: (2024)
by: Ye, Yaoqin, et al.
Published: (2024)
An Explainable Biomedical Foundation Model via Large-Scale Concept-Enhanced Vision-Language Pre-training
by: Nie, Yuxiang, et al.
Published: (2025)
by: Nie, Yuxiang, et al.
Published: (2025)
Similar Items
-
FLAIR: VLM with Fine-grained Language-informed Image Representations
by: Xiao, Rui, et al.
Published: (2024) -
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025) -
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
by: Xiao, Rui, et al.
Published: (2026) -
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024) -
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
by: Girrbach, Leander, et al.
Published: (2025)