SDoH-GPT: Using Large Language Models to Extract Social Determinants of Health (SDoH)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Consoli, Bernardo, Wu, Xizhi, Wang, Song, Zhao, Xinyu, Wang, Yanshan, Rousseau, Justin, Hartvigsen, Tom, Shen, Li, Wu, Huanmei, Peng, Yifan, Long, Qi, Chen, Tianlong, Ding, Ying
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917731775807488
author Consoli, Bernardo
Wu, Xizhi
Wang, Song
Zhao, Xinyu
Wang, Yanshan
Rousseau, Justin
Hartvigsen, Tom
Shen, Li
Wu, Huanmei
Peng, Yifan
Long, Qi
Chen, Tianlong
Ding, Ying
author_facet Consoli, Bernardo
Wu, Xizhi
Wang, Song
Zhao, Xinyu
Wang, Yanshan
Rousseau, Justin
Hartvigsen, Tom
Shen, Li
Wu, Huanmei
Peng, Yifan
Long, Qi
Chen, Tianlong
Ding, Ying
contents Extracting social determinants of health (SDoH) from unstructured medical notes depends heavily on labor-intensive annotations, which are typically task-specific, hampering reusability and limiting sharing. In this study we introduced SDoH-GPT, a simple and effective few-shot Large Language Model (LLM) method leveraging contrastive examples and concise instructions to extract SDoH without relying on extensive medical annotations or costly human intervention. It achieved tenfold and twentyfold reductions in time and cost respectively, and superior consistency with human annotators measured by Cohen's kappa of up to 0.92. The innovative combination of SDoH-GPT and XGBoost leverages the strengths of both, ensuring high accuracy and computational efficiency while consistently maintaining 0.90+ AUROC scores. Testing across three distinct datasets has confirmed its robustness and accuracy. This study highlights the potential of leveraging LLMs to revolutionize medical note classification, demonstrating their capability to achieve highly accurate classifications with significantly reduced time and cost.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17126
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SDoH-GPT: Using Large Language Models to Extract Social Determinants of Health (SDoH)
Consoli, Bernardo
Wu, Xizhi
Wang, Song
Zhao, Xinyu
Wang, Yanshan
Rousseau, Justin
Hartvigsen, Tom
Shen, Li
Wu, Huanmei
Peng, Yifan
Long, Qi
Chen, Tianlong
Ding, Ying
Computation and Language
Artificial Intelligence
Extracting social determinants of health (SDoH) from unstructured medical notes depends heavily on labor-intensive annotations, which are typically task-specific, hampering reusability and limiting sharing. In this study we introduced SDoH-GPT, a simple and effective few-shot Large Language Model (LLM) method leveraging contrastive examples and concise instructions to extract SDoH without relying on extensive medical annotations or costly human intervention. It achieved tenfold and twentyfold reductions in time and cost respectively, and superior consistency with human annotators measured by Cohen's kappa of up to 0.92. The innovative combination of SDoH-GPT and XGBoost leverages the strengths of both, ensuring high accuracy and computational efficiency while consistently maintaining 0.90+ AUROC scores. Testing across three distinct datasets has confirmed its robustness and accuracy. This study highlights the potential of leveraging LLMs to revolutionize medical note classification, demonstrating their capability to achieve highly accurate classifications with significantly reduced time and cost.
title SDoH-GPT: Using Large Language Models to Extract Social Determinants of Health (SDoH)
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.17126