Latent Covariate Shift: Unlocking Partial Identifiability for Multi-Source Domain Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuhang, Zhang, Zhen, Gong, Dong, Gong, Mingming, Huang, Biwei, Hengel, Anton van den, Zhang, Kun, Shi, Javen Qinfeng
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915222316384256
author Liu, Yuhang
Zhang, Zhen
Gong, Dong
Gong, Mingming
Huang, Biwei
Hengel, Anton van den
Zhang, Kun
Shi, Javen Qinfeng
author_facet Liu, Yuhang
Zhang, Zhen
Gong, Dong
Gong, Mingming
Huang, Biwei
Hengel, Anton van den
Zhang, Kun
Shi, Javen Qinfeng
contents Multi-source domain adaptation (MSDA) addresses the challenge of learning a label prediction function for an unlabeled target domain by leveraging both the labeled data from multiple source domains and the unlabeled data from the target domain. Conventional MSDA approaches often rely on covariate shift or conditional shift paradigms, which assume a consistent label distribution across domains. However, this assumption proves limiting in practical scenarios where label distributions do vary across domains, diminishing its applicability in real-world settings. For example, animals from different regions exhibit diverse characteristics due to varying diets and genetics. Motivated by this, we propose a novel paradigm called latent covariate shift (LCS), which introduces significantly greater variability and adaptability across domains. Notably, it provides a theoretical assurance for recovering the latent cause of the label variable, which we refer to as the latent content variable. Within this new paradigm, we present an intricate causal generative model by introducing latent noises across domains, along with a latent content variable and a latent style variable to achieve more nuanced rendering of observational data. We demonstrate that the latent content variable can be identified up to block identifiability due to its versatile yet distinct causal structure. We anchor our theoretical insights into a novel MSDA method, which learns the label distribution conditioned on the identifiable latent content variable, thereby accommodating more substantial distribution shifts. The proposed approach showcases exceptional performance and efficacy on both simulated and real-world datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2208_14161
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Latent Covariate Shift: Unlocking Partial Identifiability for Multi-Source Domain Adaptation
Liu, Yuhang
Zhang, Zhen
Gong, Dong
Gong, Mingming
Huang, Biwei
Hengel, Anton van den
Zhang, Kun
Shi, Javen Qinfeng
Machine Learning
Multi-source domain adaptation (MSDA) addresses the challenge of learning a label prediction function for an unlabeled target domain by leveraging both the labeled data from multiple source domains and the unlabeled data from the target domain. Conventional MSDA approaches often rely on covariate shift or conditional shift paradigms, which assume a consistent label distribution across domains. However, this assumption proves limiting in practical scenarios where label distributions do vary across domains, diminishing its applicability in real-world settings. For example, animals from different regions exhibit diverse characteristics due to varying diets and genetics. Motivated by this, we propose a novel paradigm called latent covariate shift (LCS), which introduces significantly greater variability and adaptability across domains. Notably, it provides a theoretical assurance for recovering the latent cause of the label variable, which we refer to as the latent content variable. Within this new paradigm, we present an intricate causal generative model by introducing latent noises across domains, along with a latent content variable and a latent style variable to achieve more nuanced rendering of observational data. We demonstrate that the latent content variable can be identified up to block identifiability due to its versatile yet distinct causal structure. We anchor our theoretical insights into a novel MSDA method, which learns the label distribution conditioned on the identifiable latent content variable, thereby accommodating more substantial distribution shifts. The proposed approach showcases exceptional performance and efficacy on both simulated and real-world datasets.
title Latent Covariate Shift: Unlocking Partial Identifiability for Multi-Source Domain Adaptation
topic Machine Learning
url https://arxiv.org/abs/2208.14161