Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hashtarkhani, Soheil, Rashid, Rezaur, Brett, Christopher L, Chinthala, Lokesh, Kumsa, Fekede Asefa, Zink, Janet A, Davis, Robert L, Schwartz, David L, Shaban-Nejad, Arash
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917014847619072
author Hashtarkhani, Soheil
Rashid, Rezaur
Brett, Christopher L
Chinthala, Lokesh
Kumsa, Fekede Asefa
Zink, Janet A
Davis, Robert L
Schwartz, David L
Shaban-Nejad, Arash
author_facet Hashtarkhani, Soheil
Rashid, Rezaur
Brett, Christopher L
Chinthala, Lokesh
Kumsa, Fekede Asefa
Zink, Janet A
Davis, Robert L
Schwartz, David L
Shaban-Nejad, Arash
contents Electronic health records contain inconsistently structured or free-text data, requiring efficient preprocessing to enable predictive health care models. Although artificial intelligence-driven natural language processing tools show promise for automating diagnosis classification, their comparative performance and clinical reliability require systematic evaluation. The aim of this study is to evaluate the performance of 4 large language models (GPT-3.5, GPT-4o, Llama 3.2, and Gemini 1.5) and BioBERT in classifying cancer diagnoses from structured and unstructured electronic health records data. We analyzed 762 unique diagnoses (326 International Classification of Diseases (ICD) code descriptions, 436free-text entries) from 3456 records of patients with cancer. Models were tested on their ability to categorize diagnoses into 14predefined categories. Two oncology experts validated classifications. BioBERT achieved the highest weighted macro F1-score for ICD codes (84.2) and matched GPT-4o in ICD code accuracy (90.8). For free-text diagnoses, GPT-4o outperformed BioBERT in weighted macro F1-score (71.8 vs 61.5) and achieved slightly higher accuracy (81.9 vs 81.6). GPT-3.5, Gemini, and Llama showed lower overall performance on both formats. Common misclassification patterns included confusion between metastasis and central nervous system tumors, as well as errors involving ambiguous or overlapping clinical terminology. Although current performance levels appear sufficient for administrative and research use, reliable clinical applications will require standardized documentation practices alongside robust human oversight for high-stakes decision-making.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12813
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study
Hashtarkhani, Soheil
Rashid, Rezaur
Brett, Christopher L
Chinthala, Lokesh
Kumsa, Fekede Asefa
Zink, Janet A
Davis, Robert L
Schwartz, David L
Shaban-Nejad, Arash
Computation and Language
Artificial Intelligence
Machine Learning
Electronic health records contain inconsistently structured or free-text data, requiring efficient preprocessing to enable predictive health care models. Although artificial intelligence-driven natural language processing tools show promise for automating diagnosis classification, their comparative performance and clinical reliability require systematic evaluation. The aim of this study is to evaluate the performance of 4 large language models (GPT-3.5, GPT-4o, Llama 3.2, and Gemini 1.5) and BioBERT in classifying cancer diagnoses from structured and unstructured electronic health records data. We analyzed 762 unique diagnoses (326 International Classification of Diseases (ICD) code descriptions, 436free-text entries) from 3456 records of patients with cancer. Models were tested on their ability to categorize diagnoses into 14predefined categories. Two oncology experts validated classifications. BioBERT achieved the highest weighted macro F1-score for ICD codes (84.2) and matched GPT-4o in ICD code accuracy (90.8). For free-text diagnoses, GPT-4o outperformed BioBERT in weighted macro F1-score (71.8 vs 61.5) and achieved slightly higher accuracy (81.9 vs 81.6). GPT-3.5, Gemini, and Llama showed lower overall performance on both formats. Common misclassification patterns included confusion between metastasis and central nervous system tumors, as well as errors involving ambiguous or overlapping clinical terminology. Although current performance levels appear sufficient for administrative and research use, reliable clinical applications will require standardized documentation practices alongside robust human oversight for high-stakes decision-making.
title Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.12813