MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yim, Wen-wai, Abacha, Asma Ben, Yu, Zixuan, Doerning, Robert, Xia, Fei, Yetisgen, Meliha
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918141586571264
author Yim, Wen-wai
Abacha, Asma Ben
Yu, Zixuan
Doerning, Robert
Xia, Fei
Yetisgen, Meliha
author_facet Yim, Wen-wai
Abacha, Asma Ben
Yu, Zixuan
Doerning, Robert
Xia, Fei
Yetisgen, Meliha
contents Evaluating natural language generation (NLG) systems in the medical domain presents unique challenges due to the critical demands for accuracy, relevance, and domain-specific expertise. Traditional automatic evaluation metrics, such as BLEU, ROUGE, and BERTScore, often fall short in distinguishing between high-quality outputs, especially given the open-ended nature of medical question answering (QA) tasks where multiple valid responses may exist. In this work, we introduce MORQA (Medical Open-Response QA), a new multilingual benchmark designed to assess the effectiveness of NLG evaluation metrics across three medical visual and text-based QA datasets in English and Chinese. Unlike prior resources, our datasets feature 2-4+ gold-standard answers authored by medical professionals, along with expert human ratings for three English and Chinese subsets. We benchmark both traditional metrics and large language model (LLM)-based evaluators, such as GPT-4 and Gemini, finding that LLM-based approaches significantly outperform traditional metrics in correlating with expert judgments. We further analyze factors driving this improvement, including LLMs' sensitivity to semantic nuances and robustness to variability among reference answers. Our results provide the first comprehensive, multilingual qualitative study of NLG evaluation in the medical domain, highlighting the need for human-aligned evaluation methods. All datasets and annotations will be publicly released to support future research.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
Yim, Wen-wai
Abacha, Asma Ben
Yu, Zixuan
Doerning, Robert
Xia, Fei
Yetisgen, Meliha
Computation and Language
68T50 (Primary) 68T45 (Secondary)
I.2.7; I.2.10
Evaluating natural language generation (NLG) systems in the medical domain presents unique challenges due to the critical demands for accuracy, relevance, and domain-specific expertise. Traditional automatic evaluation metrics, such as BLEU, ROUGE, and BERTScore, often fall short in distinguishing between high-quality outputs, especially given the open-ended nature of medical question answering (QA) tasks where multiple valid responses may exist. In this work, we introduce MORQA (Medical Open-Response QA), a new multilingual benchmark designed to assess the effectiveness of NLG evaluation metrics across three medical visual and text-based QA datasets in English and Chinese. Unlike prior resources, our datasets feature 2-4+ gold-standard answers authored by medical professionals, along with expert human ratings for three English and Chinese subsets. We benchmark both traditional metrics and large language model (LLM)-based evaluators, such as GPT-4 and Gemini, finding that LLM-based approaches significantly outperform traditional metrics in correlating with expert judgments. We further analyze factors driving this improvement, including LLMs' sensitivity to semantic nuances and robustness to variability among reference answers. Our results provide the first comprehensive, multilingual qualitative study of NLG evaluation in the medical domain, highlighting the need for human-aligned evaluation methods. All datasets and annotations will be publicly released to support future research.
title MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
topic Computation and Language
68T50 (Primary) 68T45 (Secondary)
I.2.7; I.2.10
url https://arxiv.org/abs/2509.12405