A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rogoz, Ana-Cristina, Ionescu, Radu Tudor, Anghel, Alexandra-Valentina, Antone-Iordache, Ionut-Lucian, Coniac, Simona, Ionescu, Andreea Iuliana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912898334326784
author Rogoz, Ana-Cristina
Ionescu, Radu Tudor
Anghel, Alexandra-Valentina
Antone-Iordache, Ionut-Lucian
Coniac, Simona
Ionescu, Andreea Iuliana
author_facet Rogoz, Ana-Cristina
Ionescu, Radu Tudor
Anghel, Alexandra-Valentina
Antone-Iordache, Ionut-Lucian
Coniac, Simona
Ionescu, Andreea Iuliana
contents We introduce MedQARo, the first large-scale medical QA benchmark in Romanian, alongside a comprehensive evaluation of state-of-the-art large language models (LLMs). We construct a high-quality and large-scale dataset comprising 105,880 QA pairs about cancer patients from two medical centers. The questions regard medical case summaries of 1,242 patients, requiring both keyword extraction and reasoning. Our benchmark contains both in-domain and cross-domain (cross-center and cross-cancer) test collections, enabling a precise assessment of generalization capabilities. We experiment with four open-source LLMs from distinct families of models on MedQARo. Each model is employed in two scenarios: zero-shot prompting and supervised fine-tuning. We also evaluate two state-of-the-art LLMs exposed only through APIs, namely GPT-5.2 and Gemini 3 Flash. Our results show that fine-tuned models significantly outperform zero-shot models, indicating that pretrained models fail to generalize on MedQARo. Our findings demonstrate the importance of both domain-specific and language-specific fine-tuning for reliable clinical QA in Romanian.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian
Rogoz, Ana-Cristina
Ionescu, Radu Tudor
Anghel, Alexandra-Valentina
Antone-Iordache, Ionut-Lucian
Coniac, Simona
Ionescu, Andreea Iuliana
Computation and Language
Artificial Intelligence
Machine Learning
We introduce MedQARo, the first large-scale medical QA benchmark in Romanian, alongside a comprehensive evaluation of state-of-the-art large language models (LLMs). We construct a high-quality and large-scale dataset comprising 105,880 QA pairs about cancer patients from two medical centers. The questions regard medical case summaries of 1,242 patients, requiring both keyword extraction and reasoning. Our benchmark contains both in-domain and cross-domain (cross-center and cross-cancer) test collections, enabling a precise assessment of generalization capabilities. We experiment with four open-source LLMs from distinct families of models on MedQARo. Each model is employed in two scenarios: zero-shot prompting and supervised fine-tuning. We also evaluate two state-of-the-art LLMs exposed only through APIs, namely GPT-5.2 and Gemini 3 Flash. Our results show that fine-tuned models significantly outperform zero-shot models, indicating that pretrained models fail to generalize on MedQARo. Our findings demonstrate the importance of both domain-specific and language-specific fine-tuning for reliable clinical QA in Romanian.
title A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.16390