Performance Assessment Strategies for Language Model Applications in Healthcare

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Garcia, Victor, Sidulova, Mariia, Badano, Aldo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910043113259008
author Garcia, Victor
Sidulova, Mariia
Badano, Aldo
author_facet Garcia, Victor
Sidulova, Mariia
Badano, Aldo
contents Language models (LMs) represent an emerging paradigm within artificial intelligence, with applications throughout the medical enterprise. A comprehensive understanding of the clinical task and awareness of the variability in performance when implemented in actual clinical environments lays the foundation for the LM application assessment. Presently, a prevalent method for evaluating the performance of these generative models relies on quantitative benchmarks. Such benchmarks have limitations and may suffer from train-to-the-test overfitting, optimizing performance for a specified test set at the cost of generalizability across other tasks and data distributions. Evaluation strategies leveraging human expertise and utilizing cost-effective computational models as evaluators are gaining interest. We discuss current state-of-the-art methodologies for assessing the performance of LM applications in healthcare and medical devices.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08087
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Performance Assessment Strategies for Language Model Applications in Healthcare
Garcia, Victor
Sidulova, Mariia
Badano, Aldo
Machine Learning
Artificial Intelligence
Language models (LMs) represent an emerging paradigm within artificial intelligence, with applications throughout the medical enterprise. A comprehensive understanding of the clinical task and awareness of the variability in performance when implemented in actual clinical environments lays the foundation for the LM application assessment. Presently, a prevalent method for evaluating the performance of these generative models relies on quantitative benchmarks. Such benchmarks have limitations and may suffer from train-to-the-test overfitting, optimizing performance for a specified test set at the cost of generalizability across other tasks and data distributions. Evaluation strategies leveraging human expertise and utilizing cost-effective computational models as evaluators are gaining interest. We discuss current state-of-the-art methodologies for assessing the performance of LM applications in healthcare and medical devices.
title Performance Assessment Strategies for Language Model Applications in Healthcare
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.08087