Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jang, Young Kyun, Lim, Ser-nam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language Models
by: Yang, Yongjin, et al.
Published: (2023)
by: Yang, Yongjin, et al.
Published: (2023)
ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts
by: Choi, Sangbum, et al.
Published: (2025)
by: Choi, Sangbum, et al.
Published: (2025)
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
by: He, Yuting, et al.
Published: (2026)
by: He, Yuting, et al.
Published: (2026)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
by: Zhang, Yudong, et al.
Published: (2024)
by: Zhang, Yudong, et al.
Published: (2024)
Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
by: Pramanick, Shraman, et al.
Published: (2023)
by: Pramanick, Shraman, et al.
Published: (2023)
Variational Adapter for Cross-modal Similarity Representation
by: Wei, WenZhang, et al.
Published: (2026)
by: Wei, WenZhang, et al.
Published: (2026)
MGHanD: Multi-modal Guidance for authentic Hand Diffusion
by: Eum, Taehyeon, et al.
Published: (2025)
by: Eum, Taehyeon, et al.
Published: (2025)
Diffexplainer: Towards Cross-modal Global Explanations with Diffusion Models
by: Pennisi, Matteo, et al.
Published: (2024)
by: Pennisi, Matteo, et al.
Published: (2024)
ViT-Lens: Towards Omni-modal Representations
by: Lei, Weixian, et al.
Published: (2023)
by: Lei, Weixian, et al.
Published: (2023)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
by: Walimbe, Soham, et al.
Published: (2025)
by: Walimbe, Soham, et al.
Published: (2025)
Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
by: Kim, Hayeon, et al.
Published: (2026)
by: Kim, Hayeon, et al.
Published: (2026)
Towards Calibrated Robust Fine-Tuning of Vision-Language Models
by: Oh, Changdae, et al.
Published: (2023)
by: Oh, Changdae, et al.
Published: (2023)
MATE: Meet At The Embedding -- Connecting Images with Long Texts
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
Explaining Multi-modal Large Language Models by Analyzing their Vision Perception
by: Giulivi, Loris, et al.
Published: (2024)
by: Giulivi, Loris, et al.
Published: (2024)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
by: Chen, Harold Haodong, et al.
Published: (2024)
by: Chen, Harold Haodong, et al.
Published: (2024)
Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment
by: Li, Jiaqing, et al.
Published: (2026)
by: Li, Jiaqing, et al.
Published: (2026)
Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma Grading
by: Pan, Li, et al.
Published: (2024)
by: Pan, Li, et al.
Published: (2024)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022)
by: Li, Bing, et al.
Published: (2022)
The Geometry of Representational Failures in Vision Language Models
by: Savietto, Daniele, et al.
Published: (2026)
by: Savietto, Daniele, et al.
Published: (2026)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
by: Yao, Lei, et al.
Published: (2025)
by: Yao, Lei, et al.
Published: (2025)
Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning
by: Bleeker, Maurits, et al.
Published: (2024)
by: Bleeker, Maurits, et al.
Published: (2024)
Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language Models
by: Chen, Biao, et al.
Published: (2025)
by: Chen, Biao, et al.
Published: (2025)
A Unified Debiasing Approach for Vision-Language Models across Modalities and Tasks
by: Jung, Hoin, et al.
Published: (2024)
by: Jung, Hoin, et al.
Published: (2024)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
Cross-modal Information Flow in Multimodal Large Language Models
by: Zhang, Zhi, et al.
Published: (2024)
by: Zhang, Zhi, et al.
Published: (2024)
Representation Calibration and Uncertainty Guidance for Class-Incremental Learning based on Vision Language Model
by: Tan, Jiantao, et al.
Published: (2025)
by: Tan, Jiantao, et al.
Published: (2025)
Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation
by: Tan, Zhaorui, et al.
Published: (2024)
by: Tan, Zhaorui, et al.
Published: (2024)
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
by: Meyarian, Abolfazl, et al.
Published: (2026)
by: Meyarian, Abolfazl, et al.
Published: (2026)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
by: Li, Yanwei, et al.
Published: (2024)
by: Li, Yanwei, et al.
Published: (2024)
GTMA: Dynamic Representation Optimization for OOD Vision-Language Models
by: Zhang, Jensen, et al.
Published: (2025)
by: Zhang, Jensen, et al.
Published: (2025)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
by: Lim, Hyeonseok, et al.
Published: (2024)
by: Lim, Hyeonseok, et al.
Published: (2024)
AirSketch: Generative Motion to Sketch
by: Lim, Hui Xian Grace, et al.
Published: (2024)
by: Lim, Hui Xian Grace, et al.
Published: (2024)
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
by: Xu, Dunyuan, et al.
Published: (2024)
by: Xu, Dunyuan, et al.
Published: (2024)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
by: Cheng, Zihui, et al.
Published: (2024)
by: Cheng, Zihui, et al.
Published: (2024)
Similar Items
-
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
by: Jang, Young Kyun, et al.
Published: (2024) -
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
by: Jang, Young Kyun, et al.
Published: (2024) -
Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval
by: Jang, Young Kyun, et al.
Published: (2024) -
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
by: Mohammadshirazi, Ahmad, et al.
Published: (2024) -
Towards Difficulty-Agnostic Efficient Transfer Learning for Vision-Language Models
by: Yang, Yongjin, et al.
Published: (2023)