Saved in:
Bibliographic Details
Main Authors: Wang, Linna, You, Zhixuan, Zhang, Qihui, Wen, Jiunan, Shi, Ji, Chen, Yimin, Wang, Yusen, Ding, Fanqi, Feng, Ziliang, Lu, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.07127
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918200168415232
author Wang, Linna
You, Zhixuan
Zhang, Qihui
Wen, Jiunan
Shi, Ji
Chen, Yimin
Wang, Yusen
Ding, Fanqi
Feng, Ziliang
Lu, Li
author_facet Wang, Linna
You, Zhixuan
Zhang, Qihui
Wen, Jiunan
Shi, Ji
Chen, Yimin
Wang, Yusen
Ding, Fanqi
Feng, Ziliang
Lu, Li
contents Large Language Models (LLMs) and causal learning each hold strong potential for clinical decision making (CDM). However, their synergy remains poorly understood, largely due to the lack of systematic benchmarks evaluating their integration in clinical risk prediction. In real-world healthcare, identifying features with causal influence on outcomes is crucial for actionable and trustworthy predictions. While recent work highlights LLMs' emerging causal reasoning abilities, there lacks comprehensive benchmarks to assess their causal learning and performance informed by causal features in clinical risk prediction. To address this, we introduce REACT-LLM, a benchmark designed to evaluate whether combining LLMs with causal features can enhance clinical prognostic performance and potentially outperform traditional machine learning (ML) methods. Unlike existing LLM-clinical benchmarks that often focus on a limited set of outcomes, REACT-LLM evaluates 7 clinical outcomes across 2 real-world datasets, comparing 15 prominent LLMs, 6 traditional ML models, and 3 causal discovery (CD) algorithms. Our findings indicate that while LLMs perform reasonably in clinical prognostics, they have not yet outperformed traditional ML models. Integrating causal features derived from CD algorithms into LLMs offers limited performance gains, primarily due to the strict assumptions of many CD methods, which are often violated in complex clinical data. While the direct integration yields limited improvement, our benchmark reveals a more promising synergy.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic Tasks
Wang, Linna
You, Zhixuan
Zhang, Qihui
Wen, Jiunan
Shi, Ji
Chen, Yimin
Wang, Yusen
Ding, Fanqi
Feng, Ziliang
Lu, Li
Machine Learning
Large Language Models (LLMs) and causal learning each hold strong potential for clinical decision making (CDM). However, their synergy remains poorly understood, largely due to the lack of systematic benchmarks evaluating their integration in clinical risk prediction. In real-world healthcare, identifying features with causal influence on outcomes is crucial for actionable and trustworthy predictions. While recent work highlights LLMs' emerging causal reasoning abilities, there lacks comprehensive benchmarks to assess their causal learning and performance informed by causal features in clinical risk prediction. To address this, we introduce REACT-LLM, a benchmark designed to evaluate whether combining LLMs with causal features can enhance clinical prognostic performance and potentially outperform traditional machine learning (ML) methods. Unlike existing LLM-clinical benchmarks that often focus on a limited set of outcomes, REACT-LLM evaluates 7 clinical outcomes across 2 real-world datasets, comparing 15 prominent LLMs, 6 traditional ML models, and 3 causal discovery (CD) algorithms. Our findings indicate that while LLMs perform reasonably in clinical prognostics, they have not yet outperformed traditional ML models. Integrating causal features derived from CD algorithms into LLMs offers limited performance gains, primarily due to the strict assumptions of many CD methods, which are often violated in complex clinical data. While the direct integration yields limited improvement, our benchmark reveals a more promising synergy.
title REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic Tasks
topic Machine Learning
url https://arxiv.org/abs/2511.07127