Benchmarking LLMs and SLMs for patient reported outcomes

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Marengo, Matteo, Lévy, Jarod, Bibault, Jean-Emmanuel
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917875984367616
author Marengo, Matteo
Lévy, Jarod
Bibault, Jean-Emmanuel
author_facet Marengo, Matteo
Lévy, Jarod
Bibault, Jean-Emmanuel
contents LLMs have transformed the execution of numerous tasks, including those in the medical domain. Among these, summarizing patient-reported outcomes (PROs) into concise natural language reports is of particular interest to clinicians, as it enables them to focus on critical patient concerns and spend more time in meaningful discussions. While existing work with LLMs like GPT-4 has shown impressive results, real breakthroughs could arise from leveraging SLMs as they offer the advantage of being deployable locally, ensuring patient data privacy and compliance with healthcare regulations. This study benchmarks several SLMs against LLMs for summarizing patient-reported Q\&A forms in the context of radiotherapy. Using various metrics, we evaluate their precision and reliability. The findings highlight both the promise and limitations of SLMs for high-stakes medical tasks, fostering more efficient and privacy-preserving AI-driven healthcare solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16291
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benchmarking LLMs and SLMs for patient reported outcomes
Marengo, Matteo
Lévy, Jarod
Bibault, Jean-Emmanuel
Artificial Intelligence
Computation and Language
LLMs have transformed the execution of numerous tasks, including those in the medical domain. Among these, summarizing patient-reported outcomes (PROs) into concise natural language reports is of particular interest to clinicians, as it enables them to focus on critical patient concerns and spend more time in meaningful discussions. While existing work with LLMs like GPT-4 has shown impressive results, real breakthroughs could arise from leveraging SLMs as they offer the advantage of being deployable locally, ensuring patient data privacy and compliance with healthcare regulations. This study benchmarks several SLMs against LLMs for summarizing patient-reported Q\&A forms in the context of radiotherapy. Using various metrics, we evaluate their precision and reliability. The findings highlight both the promise and limitations of SLMs for high-stakes medical tasks, fostering more efficient and privacy-preserving AI-driven healthcare solutions.
title Benchmarking LLMs and SLMs for patient reported outcomes
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.16291