A comparative study of zero-shot inference with large language models and supervised modeling in breast cancer pathology classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sushil, Madhumita, Zack, Travis, Mandair, Divneet, Zheng, Zhiwei, Wali, Ahmed, Yu, Yan-Ning, Quan, Yuwei, Butte, Atul J.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913208921489408
author Sushil, Madhumita
Zack, Travis
Mandair, Divneet
Zheng, Zhiwei
Wali, Ahmed
Yu, Yan-Ning
Quan, Yuwei
Butte, Atul J.
author_facet Sushil, Madhumita
Zack, Travis
Mandair, Divneet
Zheng, Zhiwei
Wali, Ahmed
Yu, Yan-Ning
Quan, Yuwei
Butte, Atul J.
contents Although supervised machine learning is popular for information extraction from clinical notes, creating large annotated datasets requires extensive domain expertise and is time-consuming. Meanwhile, large language models (LLMs) have demonstrated promising transfer learning capability. In this study, we explored whether recent LLMs can reduce the need for large-scale data annotations. We curated a manually-labeled dataset of 769 breast cancer pathology reports, labeled with 13 categories, to compare zero-shot classification capability of the GPT-4 model and the GPT-3.5 model with supervised classification performance of three model architectures: random forests classifier, long short-term memory networks with attention (LSTM-Att), and the UCSF-BERT model. Across all 13 tasks, the GPT-4 model performed either significantly better than or as well as the best supervised model, the LSTM-Att model (average macro F1 score of 0.83 vs. 0.75). On tasks with high imbalance between labels, the differences were more prominent. Frequent sources of GPT-4 errors included inferences from multiple samples and complex task design. On complex tasks where large annotated datasets cannot be easily collected, LLMs can reduce the burden of large-scale data labeling. However, if the use of LLMs is prohibitive, the use of simpler supervised models with large annotated datasets can provide comparable results. LLMs demonstrated the potential to speed up the execution of clinical NLP studies by reducing the need for curating large annotated datasets. This may result in an increase in the utilization of NLP-based variables and outcomes in observational clinical studies.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13887
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A comparative study of zero-shot inference with large language models and supervised modeling in breast cancer pathology classification
Sushil, Madhumita
Zack, Travis
Mandair, Divneet
Zheng, Zhiwei
Wali, Ahmed
Yu, Yan-Ning
Quan, Yuwei
Butte, Atul J.
Computation and Language
Machine Learning
Although supervised machine learning is popular for information extraction from clinical notes, creating large annotated datasets requires extensive domain expertise and is time-consuming. Meanwhile, large language models (LLMs) have demonstrated promising transfer learning capability. In this study, we explored whether recent LLMs can reduce the need for large-scale data annotations. We curated a manually-labeled dataset of 769 breast cancer pathology reports, labeled with 13 categories, to compare zero-shot classification capability of the GPT-4 model and the GPT-3.5 model with supervised classification performance of three model architectures: random forests classifier, long short-term memory networks with attention (LSTM-Att), and the UCSF-BERT model. Across all 13 tasks, the GPT-4 model performed either significantly better than or as well as the best supervised model, the LSTM-Att model (average macro F1 score of 0.83 vs. 0.75). On tasks with high imbalance between labels, the differences were more prominent. Frequent sources of GPT-4 errors included inferences from multiple samples and complex task design. On complex tasks where large annotated datasets cannot be easily collected, LLMs can reduce the burden of large-scale data labeling. However, if the use of LLMs is prohibitive, the use of simpler supervised models with large annotated datasets can provide comparable results. LLMs demonstrated the potential to speed up the execution of clinical NLP studies by reducing the need for curating large annotated datasets. This may result in an increase in the utilization of NLP-based variables and outcomes in observational clinical studies.
title A comparative study of zero-shot inference with large language models and supervised modeling in breast cancer pathology classification
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2401.13887