Revealing Treatment Non-Adherence Bias in Clinical Machine Learning Using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Zhongyuan, Suresh, Arvind, Chen, Irene Y.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917991653834752
author Liang, Zhongyuan
Suresh, Arvind
Chen, Irene Y.
author_facet Liang, Zhongyuan
Suresh, Arvind
Chen, Irene Y.
contents Machine learning systems trained on electronic health records (EHRs) increasingly guide treatment decisions, but their reliability depends on the critical assumption that patients follow the prescribed treatments recorded in EHRs. Using EHR data from 3,623 hypertension patients, we investigate how treatment non-adherence introduces implicit bias that can fundamentally distort both causal inference and predictive modeling. By extracting patient adherence information from clinical notes using a large language model (LLM), we identify 786 patients (21.7%) with medication non-adherence. We further uncover key demographic and clinical factors associated with non-adherence, as well as patient-reported reasons including side effects and difficulties obtaining refills. Our findings demonstrate that this implicit bias can not only reverse estimated treatment effects, but also degrade model performance by up to 5% while disproportionately affecting vulnerable populations by exacerbating disparities in decision outcomes and model error rates. This highlights the importance of accounting for treatment non-adherence in developing responsible and equitable clinical machine learning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revealing Treatment Non-Adherence Bias in Clinical Machine Learning Using Large Language Models
Liang, Zhongyuan
Suresh, Arvind
Chen, Irene Y.
Machine Learning
Machine learning systems trained on electronic health records (EHRs) increasingly guide treatment decisions, but their reliability depends on the critical assumption that patients follow the prescribed treatments recorded in EHRs. Using EHR data from 3,623 hypertension patients, we investigate how treatment non-adherence introduces implicit bias that can fundamentally distort both causal inference and predictive modeling. By extracting patient adherence information from clinical notes using a large language model (LLM), we identify 786 patients (21.7%) with medication non-adherence. We further uncover key demographic and clinical factors associated with non-adherence, as well as patient-reported reasons including side effects and difficulties obtaining refills. Our findings demonstrate that this implicit bias can not only reverse estimated treatment effects, but also degrade model performance by up to 5% while disproportionately affecting vulnerable populations by exacerbating disparities in decision outcomes and model error rates. This highlights the importance of accounting for treatment non-adherence in developing responsible and equitable clinical machine learning systems.
title Revealing Treatment Non-Adherence Bias in Clinical Machine Learning Using Large Language Models
topic Machine Learning
url https://arxiv.org/abs/2502.19625