Supervised Contrastive Block Disentanglement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Makino, Taro, Park, Ji Won, Tagasovska, Natasa, Kudo, Takamasa, Coelho, Paula, Huetter, Jan-Christian, Yao, Heming, Hoeckendorf, Burkhard, Leote, Ana Carolina, Ra, Stephen, Richmond, David, Cho, Kyunghyun, Regev, Aviv, Lopez, Romain
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929709863927808
author Makino, Taro
Park, Ji Won
Tagasovska, Natasa
Kudo, Takamasa
Coelho, Paula
Huetter, Jan-Christian
Yao, Heming
Hoeckendorf, Burkhard
Leote, Ana Carolina
Ra, Stephen
Richmond, David
Cho, Kyunghyun
Regev, Aviv
Lopez, Romain
author_facet Makino, Taro
Park, Ji Won
Tagasovska, Natasa
Kudo, Takamasa
Coelho, Paula
Huetter, Jan-Christian
Yao, Heming
Hoeckendorf, Burkhard
Leote, Ana Carolina
Ra, Stephen
Richmond, David
Cho, Kyunghyun
Regev, Aviv
Lopez, Romain
contents Real-world datasets often combine data collected under different experimental conditions. This yields larger datasets, but also introduces spurious correlations that make it difficult to model the phenomena of interest. We address this by learning two embeddings to independently represent the phenomena of interest and the spurious correlations. The embedding representing the phenomena of interest is correlated with the target variable $y$, and is invariant to the environment variable $e$. In contrast, the embedding representing the spurious correlations is correlated with $e$. The invariance to $e$ is difficult to achieve on real-world datasets. Our primary contribution is an algorithm called Supervised Contrastive Block Disentanglement (SCBD) that effectively enforces this invariance. It is based purely on Supervised Contrastive Learning, and applies to real-world data better than existing approaches. We empirically validate SCBD on two challenging problems. The first problem is domain generalization, where we achieve strong performance on a synthetic dataset, as well as on Camelyon17-WILDS. We introduce a single hyperparameter $α$ to control the degree of invariance to $e$. When we increase $α$ to strengthen the degree of invariance, out-of-distribution performance improves at the expense of in-distribution performance. The second problem is batch correction, in which we apply SCBD to preserve biological signal and remove inter-well batch effects when modeling single-cell perturbations from 26 million Optical Pooled Screening images.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07281
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Supervised Contrastive Block Disentanglement
Makino, Taro
Park, Ji Won
Tagasovska, Natasa
Kudo, Takamasa
Coelho, Paula
Huetter, Jan-Christian
Yao, Heming
Hoeckendorf, Burkhard
Leote, Ana Carolina
Ra, Stephen
Richmond, David
Cho, Kyunghyun
Regev, Aviv
Lopez, Romain
Machine Learning
Real-world datasets often combine data collected under different experimental conditions. This yields larger datasets, but also introduces spurious correlations that make it difficult to model the phenomena of interest. We address this by learning two embeddings to independently represent the phenomena of interest and the spurious correlations. The embedding representing the phenomena of interest is correlated with the target variable $y$, and is invariant to the environment variable $e$. In contrast, the embedding representing the spurious correlations is correlated with $e$. The invariance to $e$ is difficult to achieve on real-world datasets. Our primary contribution is an algorithm called Supervised Contrastive Block Disentanglement (SCBD) that effectively enforces this invariance. It is based purely on Supervised Contrastive Learning, and applies to real-world data better than existing approaches. We empirically validate SCBD on two challenging problems. The first problem is domain generalization, where we achieve strong performance on a synthetic dataset, as well as on Camelyon17-WILDS. We introduce a single hyperparameter $α$ to control the degree of invariance to $e$. When we increase $α$ to strengthen the degree of invariance, out-of-distribution performance improves at the expense of in-distribution performance. The second problem is batch correction, in which we apply SCBD to preserve biological signal and remove inter-well batch effects when modeling single-cell perturbations from 26 million Optical Pooled Screening images.
title Supervised Contrastive Block Disentanglement
topic Machine Learning
url https://arxiv.org/abs/2502.07281