ABC: Achieving Better Control of Multimodal Embeddings using VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Schneider, Benjamin, Kerschbaum, Florian, Chen, Wenhu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Universal Backdoor Attacks
by: Schneider, Benjamin, et al.
Published: (2023)
by: Schneider, Benjamin, et al.
Published: (2023)
Disentangling Mean Embeddings for Better Diagnostics of Image Generators
by: Gruber, Sebastian G., et al.
Published: (2024)
by: Gruber, Sebastian G., et al.
Published: (2024)
Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
by: Wang, Zhicheng, et al.
Published: (2025)
by: Wang, Zhicheng, et al.
Published: (2025)
Better Together: Evaluating the Complementarity of Earth Embedding Models
by: van der Plas, Thijs L, et al.
Published: (2026)
by: van der Plas, Thijs L, et al.
Published: (2026)
Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
by: Moon, Jihyun, et al.
Published: (2025)
by: Moon, Jihyun, et al.
Published: (2025)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models
by: Gupta, Sharut, et al.
Published: (2025)
by: Gupta, Sharut, et al.
Published: (2025)
Efficient Domain Adaptation of Multimodal Embeddings using Constrastive Learning
by: Margaritis, Georgios, et al.
Published: (2025)
by: Margaritis, Georgios, et al.
Published: (2025)
Selecting Fine-Tuning Examples by Quizzing VLMs
by: Ji, Tenghao, et al.
Published: (2025)
by: Ji, Tenghao, et al.
Published: (2025)
Exploration of VLMs for Driver Monitoring Systems Applications
by: Cañas, Paola Natalia, et al.
Published: (2025)
by: Cañas, Paola Natalia, et al.
Published: (2025)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
by: Ye, Wenqian, et al.
Published: (2024)
by: Ye, Wenqian, et al.
Published: (2024)
VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs
by: Khezresmaeilzadeh, Tina, et al.
Published: (2026)
by: Khezresmaeilzadeh, Tina, et al.
Published: (2026)
Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
Toward Inherently Robust VLMs Against Visual Perception Attacks
by: MohajerAnsari, Pedram, et al.
Published: (2025)
by: MohajerAnsari, Pedram, et al.
Published: (2025)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
by: Tekin, Selim Furkan, et al.
Published: (2026)
by: Tekin, Selim Furkan, et al.
Published: (2026)
Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs
by: Hasanebrahimi, Afsaneh, et al.
Published: (2026)
by: Hasanebrahimi, Afsaneh, et al.
Published: (2026)
Closing the Gap: Achieving Better Accuracy-Robustness Tradeoffs against Query-Based Attacks
by: Zimmer, Pascal, et al.
Published: (2023)
by: Zimmer, Pascal, et al.
Published: (2023)
CAFD: Concept-Aware DNN Fault Detection using VLMs
by: Abbasishahkoo, Amin, et al.
Published: (2026)
by: Abbasishahkoo, Amin, et al.
Published: (2026)
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
by: Xia, Canming, et al.
Published: (2026)
by: Xia, Canming, et al.
Published: (2026)
Look Closer! An Adversarial Parametric Editing Framework for Hallucination Mitigation in VLMs
by: Hu, Jiayu, et al.
Published: (2025)
by: Hu, Jiayu, et al.
Published: (2025)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025)
by: Schlarmann, Christian, et al.
Published: (2025)
Leveraging NTPs for Efficient Hallucination Detection in VLMs
by: Azachi, Ofir, et al.
Published: (2025)
by: Azachi, Ofir, et al.
Published: (2025)
Understanding and Rectifying Safety Perception Distortion in VLMs
by: Zou, Xiaohan, et al.
Published: (2025)
by: Zou, Xiaohan, et al.
Published: (2025)
Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
by: Zheng, Shunjie-Fabian, et al.
Published: (2025)
by: Zheng, Shunjie-Fabian, et al.
Published: (2025)
OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs
by: Ailuro, Stefan Maria, et al.
Published: (2026)
by: Ailuro, Stefan Maria, et al.
Published: (2026)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
by: Wang, Xiyao, et al.
Published: (2025)
by: Wang, Xiyao, et al.
Published: (2025)
T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
Benchmarking Compact VLMs for Clip-Level Surveillance Anomaly Detection Under Weak Supervision
by: Borodin, Kirill, et al.
Published: (2026)
by: Borodin, Kirill, et al.
Published: (2026)
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
by: Kolli, Govinda, et al.
Published: (2026)
by: Kolli, Govinda, et al.
Published: (2026)
Embedding-based Retrieval in Multimodal Content Moderation
by: Liang, Hanzhong, et al.
Published: (2025)
by: Liang, Hanzhong, et al.
Published: (2025)
More is Better: Deep Domain Adaptation with Multiple Sources
by: Zhao, Sicheng, et al.
Published: (2024)
by: Zhao, Sicheng, et al.
Published: (2024)
HySurvPred: Multimodal Hyperbolic Embedding with Angle-Aware Hierarchical Contrastive Learning and Uncertainty Constraints for Survival Prediction
by: Yang, Jiaqi, et al.
Published: (2025)
by: Yang, Jiaqi, et al.
Published: (2025)
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
by: Chen, Menglan, et al.
Published: (2025)
by: Chen, Menglan, et al.
Published: (2025)
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
by: Chen, Weixin, et al.
Published: (2026)
by: Chen, Weixin, et al.
Published: (2026)
GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning
by: Peng, Tiantian, et al.
Published: (2025)
by: Peng, Tiantian, et al.
Published: (2025)
From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)
by: Mishra, Suyash, et al.
Published: (2026)
by: Mishra, Suyash, et al.
Published: (2026)
Similar Items
-
Universal Backdoor Attacks
by: Schneider, Benjamin, et al.
Published: (2023) -
Disentangling Mean Embeddings for Better Diagnostics of Image Generators
by: Gruber, Sebastian G., et al.
Published: (2024) -
Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
by: Wang, Zhicheng, et al.
Published: (2025) -
Better Together: Evaluating the Complementarity of Earth Embedding Models
by: van der Plas, Thijs L, et al.
Published: (2026) -
Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
by: Moon, Jihyun, et al.
Published: (2025)