On the Domain Robustness of Contrastive Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Koddenbrock, Mario, Hoffmann, Rudolf, Brodmann, David, Rodner, Erik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feedback-driven object detection and iterative model improvement
by: Tenckhoff, Sönke, et al.
Published: (2024)
by: Tenckhoff, Sönke, et al.
Published: (2024)
LLMStructBench: Benchmarking Large Language Model Structured Data Extraction
by: Tenckhoff, Sönke, et al.
Published: (2026)
by: Tenckhoff, Sönke, et al.
Published: (2026)
Domain-Specific Self-Supervised Pre-training for Agricultural Disease Classification: A Hierarchical Vision Transformer Study
by: Sonavane, Arnav S.
Published: (2026)
by: Sonavane, Arnav S.
Published: (2026)
Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
by: Perez, Joan, et al.
Published: (2025)
by: Perez, Joan, et al.
Published: (2025)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)
by: He, Mengqi, et al.
Published: (2025)
Data-Efficient Realized Volatility Forecasting with Vision Transformers
by: Soroka, Emi, et al.
Published: (2025)
by: Soroka, Emi, et al.
Published: (2025)
SynGen-Vision: Synthetic Data Generation for training industrial vision models
by: Dubey, Alpana, et al.
Published: (2025)
by: Dubey, Alpana, et al.
Published: (2025)
Simplifying Source-Free Domain Adaptation for Object Detection: Effective Self-Training Strategies and Performance Insights
by: Hao, Yan, et al.
Published: (2024)
by: Hao, Yan, et al.
Published: (2024)
DatUS^2: Data-driven Unsupervised Semantic Segmentation with Pre-trained Self-supervised Vision Transformer
by: Kumar, Sonal, et al.
Published: (2024)
by: Kumar, Sonal, et al.
Published: (2024)
Grayscale Image Colorization with GAN and CycleGAN in Different Image Domain
by: Liang, Chen, et al.
Published: (2024)
by: Liang, Chen, et al.
Published: (2024)
A Review of Pseudo-Labeling for Computer Vision
by: Kage, Patrick, et al.
Published: (2024)
by: Kage, Patrick, et al.
Published: (2024)
Looking at Model Debiasing through the Lens of Anomaly Detection
by: Pastore, Vito Paolo, et al.
Published: (2024)
by: Pastore, Vito Paolo, et al.
Published: (2024)
VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Diffusing DeBias: Synthetic Bias Amplification for Model Debiasing
by: Ciranni, Massimiliano, et al.
Published: (2025)
by: Ciranni, Massimiliano, et al.
Published: (2025)
Lightweight Complementary-Cue Fusion for Robust Video Face Forgery Detection
by: Baek, Sunghwan, et al.
Published: (2026)
by: Baek, Sunghwan, et al.
Published: (2026)
Comparison Study: Glacier Calving Front Delineation in Synthetic Aperture Radar Images With Deep Learning
by: Gourmelon, Nora, et al.
Published: (2025)
by: Gourmelon, Nora, et al.
Published: (2025)
DIET-CP: Lightweight and Data Efficient Self Supervised Continued Pretraining
by: Rodas, Bryan, et al.
Published: (2025)
by: Rodas, Bryan, et al.
Published: (2025)
A Genealogy of Foundation Models in Remote Sensing
by: Lane, Kevin, et al.
Published: (2025)
by: Lane, Kevin, et al.
Published: (2025)
PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
by: Mueller, Felix B., et al.
Published: (2025)
by: Mueller, Felix B., et al.
Published: (2025)
General Methods Make Great Domain-specific Foundation Models: A Case-study on Fetal Ultrasound
by: Ambsdorf, Jakob, et al.
Published: (2025)
by: Ambsdorf, Jakob, et al.
Published: (2025)
PRE: Vision-Language Prompt Learning with Reparameterization Encoder
by: Pham, Thi Minh Anh, et al.
Published: (2023)
by: Pham, Thi Minh Anh, et al.
Published: (2023)
Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
by: Baxevanakis, Spiros, et al.
Published: (2026)
by: Baxevanakis, Spiros, et al.
Published: (2026)
SynthEnsemble: A Fusion of CNN, Vision Transformer, and Hybrid Models for Multi-Label Chest X-Ray Classification
by: Ashraf, S. M. Nabil, et al.
Published: (2023)
by: Ashraf, S. M. Nabil, et al.
Published: (2023)
Disentangling Generation and Regression in Stochastic Interpolants for Controllable Image Restoration
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
Boundless Across Domains: A New Paradigm of Adaptive Feature and Cross-Attention for Domain Generalization in Medical Image Segmentation
by: Xu, Yuheng, et al.
Published: (2024)
by: Xu, Yuheng, et al.
Published: (2024)
Unified Local and Global Attention Interaction Modeling for Vision Transformers
by: Nguyen, Tan, et al.
Published: (2024)
by: Nguyen, Tan, et al.
Published: (2024)
Hierarchical Pre-Training of Vision Encoders with Large Language Models
by: Lee, Eugene, et al.
Published: (2026)
by: Lee, Eugene, et al.
Published: (2026)
PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
by: Yang, Zongyou, et al.
Published: (2025)
by: Yang, Zongyou, et al.
Published: (2025)
GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
by: Hacheme, Gilles Quentin, et al.
Published: (2025)
by: Hacheme, Gilles Quentin, et al.
Published: (2025)
Uncertainty and Generalizability in Foundation Models for Earth Observation
by: Ramos-Pollan, Raul, et al.
Published: (2024)
by: Ramos-Pollan, Raul, et al.
Published: (2024)
LeDiFlow: Learned Distribution-guided Flow Matching to Accelerate Image Generation
by: Zwick, Pascal, et al.
Published: (2025)
by: Zwick, Pascal, et al.
Published: (2025)
Massively Multi-Person 3D Human Motion Forecasting with Scene Context
by: Mueller, Felix B, et al.
Published: (2024)
by: Mueller, Felix B, et al.
Published: (2024)
Tiny models from tiny data: Textual and null-text inversion for few-shot distillation
by: Landolsi, Erik, et al.
Published: (2024)
by: Landolsi, Erik, et al.
Published: (2024)
VLSlice: Interactive Vision-and-Language Slice Discovery
by: Slyman, Eric, et al.
Published: (2023)
by: Slyman, Eric, et al.
Published: (2023)
FACT: Multinomial Misalignment Classification for Point Cloud Registration
by: Dillén, Ludvig, et al.
Published: (2025)
by: Dillén, Ludvig, et al.
Published: (2025)
Guidance for Low-Level Perceptual Editing in Unconditional Diffusion Models
by: Modi, Shreyansh, et al.
Published: (2026)
by: Modi, Shreyansh, et al.
Published: (2026)
Application of Generative Adversarial Network (GAN) for Synthetic Training Data Creation to improve performance of ANN Classifier for extracting Built-Up pixels from Landsat Satellite Imagery
by: Mukherjee, Amritendu, et al.
Published: (2025)
by: Mukherjee, Amritendu, et al.
Published: (2025)
Label Delay in Online Continual Learning
by: Csaba, Botos, et al.
Published: (2023)
by: Csaba, Botos, et al.
Published: (2023)
An Immersive Multi-Elevation Multi-Seasonal Dataset for 3D Reconstruction and Visualization
by: Liu, Xijun, et al.
Published: (2024)
by: Liu, Xijun, et al.
Published: (2024)
From Misclassifications to Outliers: Joint Reliability Assessment in Classification
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
Similar Items
-
Feedback-driven object detection and iterative model improvement
by: Tenckhoff, Sönke, et al.
Published: (2024) -
LLMStructBench: Benchmarking Large Language Model Structured Data Extraction
by: Tenckhoff, Sönke, et al.
Published: (2026) -
Domain-Specific Self-Supervised Pre-training for Agricultural Disease Classification: A Hierarchical Vision Transformer Study
by: Sonavane, Arnav S.
Published: (2026) -
Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
by: Perez, Joan, et al.
Published: (2025) -
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)