Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Hao, Li, Haoxuan, Chen, Luyu, Wang, Haoxiang, Chen, Xu, Gong, Mingming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909675554865152
author Yang, Hao
Li, Haoxuan
Chen, Luyu
Wang, Haoxiang
Chen, Xu
Gong, Mingming
author_facet Yang, Hao
Li, Haoxuan
Chen, Luyu
Wang, Haoxiang
Chen, Xu
Gong, Mingming
contents Hidden confounding remains a central challenge in estimating treatment effects from observational data, as unobserved variables can lead to biased causal estimates. While recent work has explored the use of large language models (LLMs) for causal inference, most approaches still rely on the unconfoundedness assumption. In this paper, we make the first attempt to mitigate hidden confounding using LLMs. We propose ProCI (Progressive Confounder Imputation), a framework that elicits the semantic and world knowledge of LLMs to iteratively generate, impute, and validate hidden confounders. ProCI leverages two key capabilities of LLMs: their strong semantic reasoning ability, which enables the discovery of plausible confounders from both structured and unstructured inputs, and their embedded world knowledge, which supports counterfactual reasoning under latent confounding. To improve robustness, ProCI adopts a distributional reasoning strategy instead of direct value imputation to prevent the collapsed outputs. Extensive experiments demonstrate that ProCI uncovers meaningful confounders and significantly improves treatment effect estimation across various datasets and LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02928
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models
Yang, Hao
Li, Haoxuan
Chen, Luyu
Wang, Haoxiang
Chen, Xu
Gong, Mingming
Computation and Language
Artificial Intelligence
Hidden confounding remains a central challenge in estimating treatment effects from observational data, as unobserved variables can lead to biased causal estimates. While recent work has explored the use of large language models (LLMs) for causal inference, most approaches still rely on the unconfoundedness assumption. In this paper, we make the first attempt to mitigate hidden confounding using LLMs. We propose ProCI (Progressive Confounder Imputation), a framework that elicits the semantic and world knowledge of LLMs to iteratively generate, impute, and validate hidden confounders. ProCI leverages two key capabilities of LLMs: their strong semantic reasoning ability, which enables the discovery of plausible confounders from both structured and unstructured inputs, and their embedded world knowledge, which supports counterfactual reasoning under latent confounding. To improve robustness, ProCI adopts a distributional reasoning strategy instead of direct value imputation to prevent the collapsed outputs. Extensive experiments demonstrate that ProCI uncovers meaningful confounders and significantly improves treatment effect estimation across various datasets and LLMs.
title Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.02928