Fine-tuning foundational models to code diagnoses from veterinary health records

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boguslav, Mayla R., Kiehl, Adam, Kott, David, Strecker, G. Joseph, Webb, Tracy, Saklou, Nadia, Ward, Terri, Kirby, Michael
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918128057843712
author Boguslav, Mayla R.
Kiehl, Adam
Kott, David
Strecker, G. Joseph
Webb, Tracy
Saklou, Nadia
Ward, Terri
Kirby, Michael
author_facet Boguslav, Mayla R.
Kiehl, Adam
Kott, David
Strecker, G. Joseph
Webb, Tracy
Saklou, Nadia
Ward, Terri
Kirby, Michael
contents Veterinary medical records represent a large data resource for application to veterinary and One Health clinical research efforts. Use of the data is limited by interoperability challenges including inconsistent data formats and data siloing. Clinical coding using standardized medical terminologies enhances the quality of medical records and facilitates their interoperability with veterinary and human health records from other sites. Previous studies, such as DeepTag and VetTag, evaluated the application of Natural Language Processing (NLP) to automate veterinary diagnosis coding, employing long short-term memory (LSTM) and transformer models to infer a subset of Systemized Nomenclature of Medicine - Clinical Terms (SNOMED-CT) diagnosis codes from free-text clinical notes. This study expands on these efforts by incorporating all 7,739 distinct SNOMED-CT diagnosis codes recognized by the Colorado State University (CSU) Veterinary Teaching Hospital (VTH) and by leveraging the increasing availability of pre-trained language models (LMs). 13 freely-available pre-trained LMs were fine-tuned on the free-text notes from 246,473 manually-coded veterinary patient visits included in the CSU VTH's electronic health records (EHRs), which resulted in superior performance relative to previous efforts. The most accurate results were obtained when expansive labeled data were used to fine-tune relatively large clinical LMs, but the study also showed that comparable results can be obtained using more limited resources and non-clinical LMs. The results of this study contribute to the improvement of the quality of veterinary EHRs by investigating accessible methods for automated coding and support both animal and human health research by paving the way for more integrated and comprehensive health databases that span species and institutions.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15186
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-tuning foundational models to code diagnoses from veterinary health records
Boguslav, Mayla R.
Kiehl, Adam
Kott, David
Strecker, G. Joseph
Webb, Tracy
Saklou, Nadia
Ward, Terri
Kirby, Michael
Computation and Language
Artificial Intelligence
I.2.7
Veterinary medical records represent a large data resource for application to veterinary and One Health clinical research efforts. Use of the data is limited by interoperability challenges including inconsistent data formats and data siloing. Clinical coding using standardized medical terminologies enhances the quality of medical records and facilitates their interoperability with veterinary and human health records from other sites. Previous studies, such as DeepTag and VetTag, evaluated the application of Natural Language Processing (NLP) to automate veterinary diagnosis coding, employing long short-term memory (LSTM) and transformer models to infer a subset of Systemized Nomenclature of Medicine - Clinical Terms (SNOMED-CT) diagnosis codes from free-text clinical notes. This study expands on these efforts by incorporating all 7,739 distinct SNOMED-CT diagnosis codes recognized by the Colorado State University (CSU) Veterinary Teaching Hospital (VTH) and by leveraging the increasing availability of pre-trained language models (LMs). 13 freely-available pre-trained LMs were fine-tuned on the free-text notes from 246,473 manually-coded veterinary patient visits included in the CSU VTH's electronic health records (EHRs), which resulted in superior performance relative to previous efforts. The most accurate results were obtained when expansive labeled data were used to fine-tune relatively large clinical LMs, but the study also showed that comparable results can be obtained using more limited resources and non-clinical LMs. The results of this study contribute to the improvement of the quality of veterinary EHRs by investigating accessible methods for automated coding and support both animal and human health research by paving the way for more integrated and comprehensive health databases that span species and institutions.
title Fine-tuning foundational models to code diagnoses from veterinary health records
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2410.15186