Interventional Causal Representation Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ahuja, Kartik, Mahajan, Divyat, Wang, Yixin, Bengio, Yoshua
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911782254149632
author Ahuja, Kartik
Mahajan, Divyat
Wang, Yixin
Bengio, Yoshua
author_facet Ahuja, Kartik
Mahajan, Divyat
Wang, Yixin
Bengio, Yoshua
contents Causal representation learning seeks to extract high-level latent factors from low-level sensory data. Most existing methods rely on observational data and structural assumptions (e.g., conditional independence) to identify the latent factors. However, interventional data is prevalent across applications. Can interventional data facilitate causal representation learning? We explore this question in this paper. The key observation is that interventional data often carries geometric signatures of the latent factors' support (i.e. what values each latent can possibly take). For example, when the latent factors are causally connected, interventions can break the dependency between the intervened latents' support and their ancestors'. Leveraging this fact, we prove that the latent causal factors can be identified up to permutation and scaling given data from perfect $do$ interventions. Moreover, we can achieve block affine identification, namely the estimated latent factors are only entangled with a few other latents if we have access to data from imperfect interventions. These results highlight the unique power of interventional data in causal representation learning; they can enable provable identification of latent factors without any assumptions about their distributions or dependency structure.
format Preprint
id arxiv_https___arxiv_org_abs_2209_11924
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Interventional Causal Representation Learning
Ahuja, Kartik
Mahajan, Divyat
Wang, Yixin
Bengio, Yoshua
Machine Learning
Causal representation learning seeks to extract high-level latent factors from low-level sensory data. Most existing methods rely on observational data and structural assumptions (e.g., conditional independence) to identify the latent factors. However, interventional data is prevalent across applications. Can interventional data facilitate causal representation learning? We explore this question in this paper. The key observation is that interventional data often carries geometric signatures of the latent factors' support (i.e. what values each latent can possibly take). For example, when the latent factors are causally connected, interventions can break the dependency between the intervened latents' support and their ancestors'. Leveraging this fact, we prove that the latent causal factors can be identified up to permutation and scaling given data from perfect $do$ interventions. Moreover, we can achieve block affine identification, namely the estimated latent factors are only entangled with a few other latents if we have access to data from imperfect interventions. These results highlight the unique power of interventional data in causal representation learning; they can enable provable identification of latent factors without any assumptions about their distributions or dependency structure.
title Interventional Causal Representation Learning
topic Machine Learning
url https://arxiv.org/abs/2209.11924