Improving ARDS Diagnosis Through Context-Aware Concept Bottleneck Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Narain, Anish, Majumdar, Ritam, Narayanan, Nikita, Marshall, Dominic, Parbhoo, Sonali
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909735809187840
author Narain, Anish
Majumdar, Ritam
Narayanan, Nikita
Marshall, Dominic
Parbhoo, Sonali
author_facet Narain, Anish
Majumdar, Ritam
Narayanan, Nikita
Marshall, Dominic
Parbhoo, Sonali
contents Large, publicly available clinical datasets have emerged as a novel resource for understanding disease heterogeneity and to explore personalization of therapy. These datasets are derived from data not originally collected for research purposes and, as a result, are often incomplete and lack critical labels. Many AI tools have been developed to retrospectively label these datasets, such as by performing disease classification; however, they often suffer from limited interpretability. Previous work has attempted to explain predictions using Concept Bottleneck Models (CBMs), which learn interpretable concepts that map to higher-level clinical ideas, facilitating human evaluation. However, these models often experience performance limitations when the concepts fail to adequately explain or characterize the task. We use the identification of Acute Respiratory Distress Syndrome (ARDS) as a challenging test case to demonstrate the value of incorporating contextual information from clinical notes to improve CBM performance. Our approach leverages a Large Language Model (LLM) to process clinical notes and generate additional concepts, resulting in a 10% performance gain over existing methods. Additionally, it facilitates the learning of more comprehensive concepts, thereby reducing the risk of information leakage and reliance on spurious shortcuts, thus improving the characterization of ARDS.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09719
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving ARDS Diagnosis Through Context-Aware Concept Bottleneck Models
Narain, Anish
Majumdar, Ritam
Narayanan, Nikita
Marshall, Dominic
Parbhoo, Sonali
Machine Learning
Artificial Intelligence
Large, publicly available clinical datasets have emerged as a novel resource for understanding disease heterogeneity and to explore personalization of therapy. These datasets are derived from data not originally collected for research purposes and, as a result, are often incomplete and lack critical labels. Many AI tools have been developed to retrospectively label these datasets, such as by performing disease classification; however, they often suffer from limited interpretability. Previous work has attempted to explain predictions using Concept Bottleneck Models (CBMs), which learn interpretable concepts that map to higher-level clinical ideas, facilitating human evaluation. However, these models often experience performance limitations when the concepts fail to adequately explain or characterize the task. We use the identification of Acute Respiratory Distress Syndrome (ARDS) as a challenging test case to demonstrate the value of incorporating contextual information from clinical notes to improve CBM performance. Our approach leverages a Large Language Model (LLM) to process clinical notes and generate additional concepts, resulting in a 10% performance gain over existing methods. Additionally, it facilitates the learning of more comprehensive concepts, thereby reducing the risk of information leakage and reliance on spurious shortcuts, thus improving the characterization of ARDS.
title Improving ARDS Diagnosis Through Context-Aware Concept Bottleneck Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.09719