When and How Does CLIP Enable Domain and Compositional Generalization?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kempf, Elias, Schrodi, Simon, Argus, Max, Brox, Thomas
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914032867344384
author Kempf, Elias
Schrodi, Simon
Argus, Max
Brox, Thomas
author_facet Kempf, Elias
Schrodi, Simon
Argus, Max
Brox, Thomas
contents The remarkable generalization performance of contrastive vision-language models like CLIP is often attributed to the diversity of their training distributions. However, key questions remain unanswered: Can CLIP generalize to an entirely unseen domain when trained on a diverse mixture of domains (domain generalization)? Can it generalize to unseen classes within partially seen domains (compositional generalization)? What factors affect such generalization? To answer these questions, we trained CLIP models on systematically constructed training distributions with controlled domain diversity and object class exposure. Our experiments show that domain diversity is essential for both domain and compositional generalization, yet compositional generalization can be surprisingly weaker than domain generalization when the training distribution contains a suboptimal subset of the test domain. Through data-centric and mechanistic analyses, we find that successful generalization requires the learning of sufficiently shared representations in intermediate layers and circuits.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When and How Does CLIP Enable Domain and Compositional Generalization?
Kempf, Elias
Schrodi, Simon
Argus, Max
Brox, Thomas
Machine Learning
Computer Vision and Pattern Recognition
The remarkable generalization performance of contrastive vision-language models like CLIP is often attributed to the diversity of their training distributions. However, key questions remain unanswered: Can CLIP generalize to an entirely unseen domain when trained on a diverse mixture of domains (domain generalization)? Can it generalize to unseen classes within partially seen domains (compositional generalization)? What factors affect such generalization? To answer these questions, we trained CLIP models on systematically constructed training distributions with controlled domain diversity and object class exposure. Our experiments show that domain diversity is essential for both domain and compositional generalization, yet compositional generalization can be surprisingly weaker than domain generalization when the training distribution contains a suboptimal subset of the test domain. Through data-centric and mechanistic analyses, we find that successful generalization requires the learning of sufficiently shared representations in intermediate layers and circuits.
title When and How Does CLIP Enable Domain and Compositional Generalization?
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.09507