High-Throughput Phenotyping of Clinical Text Using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hier, Daniel B., Munzir, S. Ilyas, Stahlfeld, Anne, Obafemi-Ajayi, Tayo, Carrithers, Michael D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908401004445696
author Hier, Daniel B.
Munzir, S. Ilyas
Stahlfeld, Anne
Obafemi-Ajayi, Tayo
Carrithers, Michael D.
author_facet Hier, Daniel B.
Munzir, S. Ilyas
Stahlfeld, Anne
Obafemi-Ajayi, Tayo
Carrithers, Michael D.
contents High-throughput phenotyping automates the mapping of patient signs to standardized ontology concepts and is essential for precision medicine. This study evaluates the automation of phenotyping of clinical summaries from the Online Mendelian Inheritance in Man (OMIM) database using large language models. Due to their rich phenotype data, these summaries can be surrogates for physician notes. We conduct a performance comparison of GPT-4 and GPT-3.5-Turbo. Our results indicate that GPT-4 surpasses GPT-3.5-Turbo in identifying, categorizing, and normalizing signs, achieving concordance with manual annotators comparable to inter-rater agreement. Despite some limitations in sign normalization, the extensive pre-training of GPT-4 results in high performance and generalizability across several phenotyping tasks while obviating the need for manually annotated training data. Large language models are expected to be the dominant method for automating high-throughput phenotyping of clinical text.
format Preprint
id arxiv_https___arxiv_org_abs_2408_01214
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle High-Throughput Phenotyping of Clinical Text Using Large Language Models
Hier, Daniel B.
Munzir, S. Ilyas
Stahlfeld, Anne
Obafemi-Ajayi, Tayo
Carrithers, Michael D.
Computation and Language
Artificial Intelligence
I.7; I.2
High-throughput phenotyping automates the mapping of patient signs to standardized ontology concepts and is essential for precision medicine. This study evaluates the automation of phenotyping of clinical summaries from the Online Mendelian Inheritance in Man (OMIM) database using large language models. Due to their rich phenotype data, these summaries can be surrogates for physician notes. We conduct a performance comparison of GPT-4 and GPT-3.5-Turbo. Our results indicate that GPT-4 surpasses GPT-3.5-Turbo in identifying, categorizing, and normalizing signs, achieving concordance with manual annotators comparable to inter-rater agreement. Despite some limitations in sign normalization, the extensive pre-training of GPT-4 results in high performance and generalizability across several phenotyping tasks while obviating the need for manually annotated training data. Large language models are expected to be the dominant method for automating high-throughput phenotyping of clinical text.
title High-Throughput Phenotyping of Clinical Text Using Large Language Models
topic Computation and Language
Artificial Intelligence
I.7; I.2
url https://arxiv.org/abs/2408.01214