MIRA: A Bilingual Benchmark for Medical Information Response Audit

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Mengyu, Yang, Qiaoxin, Wang, Qianqian, Dai, Xiwei, Wu, Weiyi, Gao, Chongyang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913166834794496
author Xu, Mengyu
Yang, Qiaoxin
Wang, Qianqian
Dai, Xiwei
Wu, Weiyi
Gao, Chongyang
author_facet Xu, Mengyu
Yang, Qiaoxin
Wang, Qianqian
Dai, Xiwei
Wu, Weiyi
Gao, Chongyang
contents Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent judgment. We term this pattern Differential Information Dilution (DID). Language effects are model-specific rather than uniformly worse for non-English prompts. A comparison with 300 real-world health queries provides preliminary evidence of rank-order validity. A knowledge-guided mitigation prompt reduces information dilution for most models, with the largest reductions in underinformative simplification observed for Claude (~8%) and Qwen (~6%).
format Preprint
id arxiv_https___arxiv_org_abs_2605_28025
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MIRA: A Bilingual Benchmark for Medical Information Response Audit
Xu, Mengyu
Yang, Qiaoxin
Wang, Qianqian
Dai, Xiwei
Wu, Weiyi
Gao, Chongyang
Artificial Intelligence
Computation and Language
Computers and Society
Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent judgment. We term this pattern Differential Information Dilution (DID). Language effects are model-specific rather than uniformly worse for non-English prompts. A comparison with 300 real-world health queries provides preliminary evidence of rank-order validity. A knowledge-guided mitigation prompt reduces information dilution for most models, with the largest reductions in underinformative simplification observed for Claude (~8%) and Qwen (~6%).
title MIRA: A Bilingual Benchmark for Medical Information Response Audit
topic Artificial Intelligence
Computation and Language
Computers and Society
url https://arxiv.org/abs/2605.28025