Eliciting associations between clinical variables from LLMs via comparison questions across populations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kabus, Fabian, Kordtomeikel, Kian, Brox, Thomas, Wiendl, Heinz, Stolz, Daiana, Binder, Harald
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910198513270784
author Kabus, Fabian
Kordtomeikel, Kian
Brox, Thomas
Wiendl, Heinz
Stolz, Daiana
Binder, Harald
author_facet Kabus, Fabian
Kordtomeikel, Kian
Brox, Thomas
Wiendl, Heinz
Stolz, Daiana
Binder, Harald
contents The training data of large language models (LLMs) comprises a wide range of biomedical literature, reflecting data from many different patient populations. We investigate how it might be possible to recover information on correlation and causal links between patient characteristics, as a key building block for medical decision making. To avoid the pitfalls of direct elicitation, we propose an approach based on structured comparison questions, specifically patient comparison triplet questions. This is combined with a statistical model for the LLM representation that provides estimates of correlations without access to activations or model internals. Intuitively, we consider how similarity decisions of LLMs based on a first variable are affected by providing information on a second variable for one of the patients being assessed. We then induce prompt-level environment shifts to obtain correlation estimates for different subpopulations, which enables an invariant causal prediction (ICP) approach to obtain conservative candidate parent links. We demonstrate the method in two clinical domains, chronic obstructive pulmonary disease (COPD) and multiple sclerosis (MS). Across prompted environments, the elicited correlations are smooth, stable, and clinically interpretable, yet vary in a statistically significant way that supports downstream invariance testing, such that ICP provides a small set of candidate invariant parent links. These results show that indirect elicitation via triplet comparisons can recover meaningful association structure from LLMs and offer a cautious route from implicit correlations to causal statements that are congruent with LLM answering patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06335
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Eliciting associations between clinical variables from LLMs via comparison questions across populations
Kabus, Fabian
Kordtomeikel, Kian
Brox, Thomas
Wiendl, Heinz
Stolz, Daiana
Binder, Harald
Machine Learning
The training data of large language models (LLMs) comprises a wide range of biomedical literature, reflecting data from many different patient populations. We investigate how it might be possible to recover information on correlation and causal links between patient characteristics, as a key building block for medical decision making. To avoid the pitfalls of direct elicitation, we propose an approach based on structured comparison questions, specifically patient comparison triplet questions. This is combined with a statistical model for the LLM representation that provides estimates of correlations without access to activations or model internals. Intuitively, we consider how similarity decisions of LLMs based on a first variable are affected by providing information on a second variable for one of the patients being assessed. We then induce prompt-level environment shifts to obtain correlation estimates for different subpopulations, which enables an invariant causal prediction (ICP) approach to obtain conservative candidate parent links. We demonstrate the method in two clinical domains, chronic obstructive pulmonary disease (COPD) and multiple sclerosis (MS). Across prompted environments, the elicited correlations are smooth, stable, and clinically interpretable, yet vary in a statistically significant way that supports downstream invariance testing, such that ICP provides a small set of candidate invariant parent links. These results show that indirect elicitation via triplet comparisons can recover meaningful association structure from LLMs and offer a cautious route from implicit correlations to causal statements that are congruent with LLM answering patterns.
title Eliciting associations between clinical variables from LLMs via comparison questions across populations
topic Machine Learning
url https://arxiv.org/abs/2605.06335