DeCode: Decoupling Content and Delivery for Medical QA

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ko, Po-Jen, Tsai, Chen-Han, Peng, Yu-Shao
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911510861709312
author Ko, Po-Jen
Tsai, Chen-Han
Peng, Yu-Shao
author_facet Ko, Po-Jen
Tsai, Chen-Han
Peng, Yu-Shao
contents Large language models (LLMs) exhibit strong medical knowledge and can generate factually accurate responses. However, existing models often fail to account for individual patient contexts, producing answers that are clinically correct yet poorly aligned with patients' needs. In this work, we introduce DeCode (Decoupling Content and Delivery), a training-free, model-agnostic framework that adapts existing LLMs to produce contextualized answers in clinical settings. We evaluate DeCode on OpenAI HealthBench, a comprehensive and challenging benchmark designed to assess clinical relevance and validity of LLM responses. DeCode boosts zero-shot performance from 28.4% to 49.8% and achieves new state-of-the-art compared to existing methods. Experimental results suggest the effectiveness of DeCode in improving clinical question answering of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02123
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DeCode: Decoupling Content and Delivery for Medical QA
Ko, Po-Jen
Tsai, Chen-Han
Peng, Yu-Shao
Computation and Language
Artificial Intelligence
Large language models (LLMs) exhibit strong medical knowledge and can generate factually accurate responses. However, existing models often fail to account for individual patient contexts, producing answers that are clinically correct yet poorly aligned with patients' needs. In this work, we introduce DeCode (Decoupling Content and Delivery), a training-free, model-agnostic framework that adapts existing LLMs to produce contextualized answers in clinical settings. We evaluate DeCode on OpenAI HealthBench, a comprehensive and challenging benchmark designed to assess clinical relevance and validity of LLM responses. DeCode boosts zero-shot performance from 28.4% to 49.8% and achieves new state-of-the-art compared to existing methods. Experimental results suggest the effectiveness of DeCode in improving clinical question answering of LLMs.
title DeCode: Decoupling Content and Delivery for Medical QA
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.02123