Prompting language influences diagnostic reasoning and accuracy of large language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bazoge, Adrien, Corvellec, Josselin, Sid-Ahmed, Sofiane Djillali, Gourraud, Pierre-Antoine
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916025340002304
author Bazoge, Adrien
Corvellec, Josselin
Sid-Ahmed, Sofiane Djillali
Gourraud, Pierre-Antoine
author_facet Bazoge, Adrien
Corvellec, Josselin
Sid-Ahmed, Sofiane Djillali
Gourraud, Pierre-Antoine
contents Large language models (LLMs) are increasingly explored for clinical decision support, yet most evaluations are conducted in English, leaving their reliability in other languages uncertain. Here we evaluate the impact of prompting language on diagnostic reasoning and final diagnosis accuracy by comparing English and French performance across five LLMs (o3, DeepSeek-R1, GPT-4-Turbo, Llama-3.1-405B-Instruct, and BioMistral-7B). A total of 180 clinical vignettes covering 16 medical specialties were assessed by two physicians using an 18-point scale evaluating both diagnosis accuracy and reasoning quality. Four of the five models performed better in English (mean difference 0.37-0.91, adjusted p < 0.05), with the gap spanning multiple aspects of reasoning, including differential diagnosis, logical structure, and internal validity. o3 was the only model showing no overall language effect. These findings demonstrate that prompting language remains a critical determinant of LLM clinical performance, with implications for equitable linguistico-cultural deployment worldwide.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19173
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Prompting language influences diagnostic reasoning and accuracy of large language models
Bazoge, Adrien
Corvellec, Josselin
Sid-Ahmed, Sofiane Djillali
Gourraud, Pierre-Antoine
Computation and Language
Large language models (LLMs) are increasingly explored for clinical decision support, yet most evaluations are conducted in English, leaving their reliability in other languages uncertain. Here we evaluate the impact of prompting language on diagnostic reasoning and final diagnosis accuracy by comparing English and French performance across five LLMs (o3, DeepSeek-R1, GPT-4-Turbo, Llama-3.1-405B-Instruct, and BioMistral-7B). A total of 180 clinical vignettes covering 16 medical specialties were assessed by two physicians using an 18-point scale evaluating both diagnosis accuracy and reasoning quality. Four of the five models performed better in English (mean difference 0.37-0.91, adjusted p < 0.05), with the gap spanning multiple aspects of reasoning, including differential diagnosis, logical structure, and internal validity. o3 was the only model showing no overall language effect. These findings demonstrate that prompting language remains a critical determinant of LLM clinical performance, with implications for equitable linguistico-cultural deployment worldwide.
title Prompting language influences diagnostic reasoning and accuracy of large language models
topic Computation and Language
url https://arxiv.org/abs/2605.19173