Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kocbek, Primoz, Kopitar, Leon, Stiglic, Gregor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909968845766656
author Kocbek, Primoz
Kopitar, Leon
Stiglic, Gregor
author_facet Kocbek, Primoz
Kopitar, Leon
Stiglic, Gregor
contents This study investigated the application of Large Language Models (LLMs) for simplifying biomedical texts to enhance health literacy. Using a public dataset, which included plain language adaptations of biomedical abstracts, we developed and evaluated several approaches, specifically a baseline approach using a prompt template, a two AI agent approach, and a fine-tuning approach. We selected OpenAI gpt-4o and gpt-4o mini models as baselines for further research. We evaluated our approaches with quantitative metrics, such as Flesch-Kincaid grade level, SMOG Index, SARI, and BERTScore, G-Eval, as well as with qualitative metric, more precisely 5-point Likert scales for simplicity, accuracy, completeness, brevity. Results showed a superior performance of gpt-4o-mini and an underperformance of FT approaches. G-Eval, a LLM based quantitative metric, showed promising results, ranking the approaches similarly as the qualitative metric.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16530
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics
Kocbek, Primoz
Kopitar, Leon
Stiglic, Gregor
Computation and Language
Artificial Intelligence
I.2.7
This study investigated the application of Large Language Models (LLMs) for simplifying biomedical texts to enhance health literacy. Using a public dataset, which included plain language adaptations of biomedical abstracts, we developed and evaluated several approaches, specifically a baseline approach using a prompt template, a two AI agent approach, and a fine-tuning approach. We selected OpenAI gpt-4o and gpt-4o mini models as baselines for further research. We evaluated our approaches with quantitative metrics, such as Flesch-Kincaid grade level, SMOG Index, SARI, and BERTScore, G-Eval, as well as with qualitative metric, more precisely 5-point Likert scales for simplicity, accuracy, completeness, brevity. Results showed a superior performance of gpt-4o-mini and an underperformance of FT approaches. G-Eval, a LLM based quantitative metric, showed promising results, ranking the approaches similarly as the qualitative metric.
title Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2512.16530