BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jonker, Richard A. A., Christiansen, Alexander, Maniatis, Alexandros, Garrido, Rúben, Lima, Rogério Braunschweiger de Freitas, Jurowetzki, Roman, Matos, Sérgio
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917462108274688
author Jonker, Richard A. A.
Christiansen, Alexander
Maniatis, Alexandros
Garrido, Rúben
Lima, Rogério Braunschweiger de Freitas
Jurowetzki, Roman
Matos, Sérgio
author_facet Jonker, Richard A. A.
Christiansen, Alexander
Maniatis, Alexandros
Garrido, Rúben
Lima, Rogério Braunschweiger de Freitas
Jurowetzki, Roman
Matos, Sérgio
contents This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without weight updates. We evaluate several state-of-the-art proprietary models and locally deployable open-source alternatives using various prompt engineering strategies, including task decomposition, Chain-of-Thought, and in-context learning. Furthermore, we explore majority voting and LLM-as-a-judge ensembling techniques to maximize predictive robustness. Our results demonstrate that while proprietary models exhibit strong resilience to prompt variations, domain-adapted open-source models (such as MedGemma 3 27B) achieve highly competitive performance when paired with the right prompt. Overall, our prompt-based approach proved highly effective, securing 1st place in Subtask 4 (evidence citation alignment) and 3rd place in Subtask 3 (patient-friendly answer generation). All code, results, and prompts are available on our GitHub repository: https://github.com/bioinformatics-ua/ArchEHR-QA-2026.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03618
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
Jonker, Richard A. A.
Christiansen, Alexander
Maniatis, Alexandros
Garrido, Rúben
Lima, Rogério Braunschweiger de Freitas
Jurowetzki, Roman
Matos, Sérgio
Computation and Language
This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without weight updates. We evaluate several state-of-the-art proprietary models and locally deployable open-source alternatives using various prompt engineering strategies, including task decomposition, Chain-of-Thought, and in-context learning. Furthermore, we explore majority voting and LLM-as-a-judge ensembling techniques to maximize predictive robustness. Our results demonstrate that while proprietary models exhibit strong resilience to prompt variations, domain-adapted open-source models (such as MedGemma 3 27B) achieve highly competitive performance when paired with the right prompt. Overall, our prompt-based approach proved highly effective, securing 1st place in Subtask 4 (evidence citation alignment) and 3rd place in Subtask 3 (patient-friendly answer generation). All code, results, and prompts are available on our GitHub repository: https://github.com/bioinformatics-ua/ArchEHR-QA-2026.
title BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
topic Computation and Language
url https://arxiv.org/abs/2605.03618