Learning Invariant Causal Mechanism from Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Zeen, Zhao, Siyu, Zhang, Xingyu, Li, Jiangmeng, Zheng, Changwen, Qiang, Wenwen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908409495814144
author Song, Zeen
Zhao, Siyu
Zhang, Xingyu
Li, Jiangmeng
Zheng, Changwen
Qiang, Wenwen
author_facet Song, Zeen
Zhao, Siyu
Zhang, Xingyu
Li, Jiangmeng
Zheng, Changwen
Qiang, Wenwen
contents Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, but its performance can degrade when fine-tuned in out-of-distribution (OOD) scenarios. We model the prediction process using a Structural Causal Model (SCM) and show that the causal mechanism involving both invariant and variant factors in training environments differs from that in test environments. In contrast, the causal mechanism with solely invariant factors remains consistent across environments. We theoretically prove the existence of a linear mapping from CLIP embeddings to invariant factors, which can be estimated using interventional data. Additionally, we provide a condition to guarantee low OOD risk of the invariant predictor. Based on these insights, we propose the Invariant Causal Mechanism of CLIP (CLIP-ICM) framework. CLIP-ICM involves collecting interventional data, estimating a linear projection matrix, and making predictions within the invariant subspace. Experiments on several OOD datasets show that CLIP-ICM significantly improves the performance of CLIP. Our method offers a simple but powerful enhancement, boosting the reliability of CLIP in real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15289
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Invariant Causal Mechanism from Vision-Language Models
Song, Zeen
Zhao, Siyu
Zhang, Xingyu
Li, Jiangmeng
Zheng, Changwen
Qiang, Wenwen
Computer Vision and Pattern Recognition
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, but its performance can degrade when fine-tuned in out-of-distribution (OOD) scenarios. We model the prediction process using a Structural Causal Model (SCM) and show that the causal mechanism involving both invariant and variant factors in training environments differs from that in test environments. In contrast, the causal mechanism with solely invariant factors remains consistent across environments. We theoretically prove the existence of a linear mapping from CLIP embeddings to invariant factors, which can be estimated using interventional data. Additionally, we provide a condition to guarantee low OOD risk of the invariant predictor. Based on these insights, we propose the Invariant Causal Mechanism of CLIP (CLIP-ICM) framework. CLIP-ICM involves collecting interventional data, estimating a linear projection matrix, and making predictions within the invariant subspace. Experiments on several OOD datasets show that CLIP-ICM significantly improves the performance of CLIP. Our method offers a simple but powerful enhancement, boosting the reliability of CLIP in real-world applications.
title Learning Invariant Causal Mechanism from Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15289