On the Domain Robustness of Contrastive Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koddenbrock, Mario, Hoffmann, Rudolf, Brodmann, David, Rodner, Erik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908428151029760
author Koddenbrock, Mario
Hoffmann, Rudolf
Brodmann, David
Rodner, Erik
author_facet Koddenbrock, Mario
Hoffmann, Rudolf
Brodmann, David
Rodner, Erik
contents In real-world vision-language applications, practitioners increasingly rely on large, pretrained foundation models rather than custom-built solutions, despite limited transparency regarding their training data and processes. While these models achieve impressive performance on general benchmarks, their effectiveness can decline notably under specialized domain shifts, such as unique imaging conditions or environmental variations. In this work, we introduce Deepbench, a framework designed to assess domain-specific robustness of vision-language models (VLMs). Deepbench leverages a large language model (LLM) to generate realistic, context-aware image corruptions tailored to specific deployment domains without requiring labeled data. We evaluate a range of contrastive vision-language architectures and architectural variants across six real-world domains and observe substantial variability in robustness, highlighting the need for targeted, domain-aware evaluation. Deepbench is released as open-source software to support further research into domain-aware robustness assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23663
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the Domain Robustness of Contrastive Vision-Language Models
Koddenbrock, Mario
Hoffmann, Rudolf
Brodmann, David
Rodner, Erik
Computer Vision and Pattern Recognition
Machine Learning
I.4
In real-world vision-language applications, practitioners increasingly rely on large, pretrained foundation models rather than custom-built solutions, despite limited transparency regarding their training data and processes. While these models achieve impressive performance on general benchmarks, their effectiveness can decline notably under specialized domain shifts, such as unique imaging conditions or environmental variations. In this work, we introduce Deepbench, a framework designed to assess domain-specific robustness of vision-language models (VLMs). Deepbench leverages a large language model (LLM) to generate realistic, context-aware image corruptions tailored to specific deployment domains without requiring labeled data. We evaluate a range of contrastive vision-language architectures and architectural variants across six real-world domains and observe substantial variability in robustness, highlighting the need for targeted, domain-aware evaluation. Deepbench is released as open-source software to support further research into domain-aware robustness assessment.
title On the Domain Robustness of Contrastive Vision-Language Models
topic Computer Vision and Pattern Recognition
Machine Learning
I.4
url https://arxiv.org/abs/2506.23663