Causal Inference for Latent Outcomes Learned with Factor Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Landy, Jenna M., Zorzetto, Dafne, De Vito, Roberta, Parmigiani, Giovanni
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915361027260416
author Landy, Jenna M.
Zorzetto, Dafne
De Vito, Roberta
Parmigiani, Giovanni
author_facet Landy, Jenna M.
Zorzetto, Dafne
De Vito, Roberta
Parmigiani, Giovanni
contents In many fields$\unicode{x2013}$including genomics, epidemiology, natural language processing, social and behavioral sciences, and economics$\unicode{x2013}$it is increasingly important to address causal questions in the context of factor models or representation learning. In this work, we investigate causal effects on $\textit{latent outcomes}$ derived from high-dimensional observed data using nonnegative matrix factorization. To the best of our knowledge, this is the first study to formally address causal inference in this setting. A central challenge is that estimating a latent factor model can cause an individual's learned latent outcome to depend on other individuals' treatments, thereby violating the standard causal inference assumption of no interference. We formalize this issue as $\textit{learning-induced interference}$ and distinguish it from interference present in a data-generating process. To address this, we propose a novel, intuitive, and theoretically grounded algorithm to estimate causal effects on latent outcomes while mitigating learning-induced interference and improving estimation efficiency. We establish theoretical guarantees for the consistency of our estimator and demonstrate its practical utility through simulation studies and an application to cancer mutational signature analysis. All baseline and proposed methods are available in our open-source R package, ${\tt causalLFO}$.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Causal Inference for Latent Outcomes Learned with Factor Models
Landy, Jenna M.
Zorzetto, Dafne
De Vito, Roberta
Parmigiani, Giovanni
Methodology
In many fields$\unicode{x2013}$including genomics, epidemiology, natural language processing, social and behavioral sciences, and economics$\unicode{x2013}$it is increasingly important to address causal questions in the context of factor models or representation learning. In this work, we investigate causal effects on $\textit{latent outcomes}$ derived from high-dimensional observed data using nonnegative matrix factorization. To the best of our knowledge, this is the first study to formally address causal inference in this setting. A central challenge is that estimating a latent factor model can cause an individual's learned latent outcome to depend on other individuals' treatments, thereby violating the standard causal inference assumption of no interference. We formalize this issue as $\textit{learning-induced interference}$ and distinguish it from interference present in a data-generating process. To address this, we propose a novel, intuitive, and theoretically grounded algorithm to estimate causal effects on latent outcomes while mitigating learning-induced interference and improving estimation efficiency. We establish theoretical guarantees for the consistency of our estimator and demonstrate its practical utility through simulation studies and an application to cancer mutational signature analysis. All baseline and proposed methods are available in our open-source R package, ${\tt causalLFO}$.
title Causal Inference for Latent Outcomes Learned with Factor Models
topic Methodology
url https://arxiv.org/abs/2506.20549