RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Jiahao, Xu, Luxin, Tan, Minghuan, Zhang, Lichao, Argha, Ahmadreza, Alinejad-Rokny, Hamid, Yang, Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909890153283584
author Zhao, Jiahao
Xu, Luxin
Tan, Minghuan
Zhang, Lichao
Argha, Ahmadreza
Alinejad-Rokny, Hamid
Yang, Min
author_facet Zhao, Jiahao
Xu, Luxin
Tan, Minghuan
Zhang, Lichao
Argha, Ahmadreza
Alinejad-Rokny, Hamid
Yang, Min
contents Numerous medical systems powered by Large Language Models (LLMs) have achieved remarkable progress in diverse healthcare tasks. However, research on their medication safety remains limited due to the lack of real world datasets, constrained by privacy and accessibility issues. Moreover, evaluation of LLMs in realistic clinical consultation settings, particularly regarding medication safety, is still underexplored. To address these gaps, we propose a framework that simulates and evaluates clinical consultations to systematically assess the medication safety capabilities of LLMs. Within this framework, we generate inquiry diagnosis dialogues with embedded medication risks and construct a dedicated medication safety database, RxRisk DB, containing 6,725 contraindications, 28,781 drug interactions, and 14,906 indication-drug pairs. A two-stage filtering strategy ensures clinical realism and professional quality, resulting in the benchmark RxSafeBench with 2,443 high-quality consultation scenarios. We evaluate leading open-source and proprietary LLMs using structured multiple choice questions that test their ability to recommend safe medications under simulated patient contexts. Results show that current LLMs struggle to integrate contraindication and interaction knowledge, especially when risks are implied rather than explicit. Our findings highlight key challenges in ensuring medication safety in LLM-based systems and provide insights into improving reliability through better prompting and task-specific tuning. RxSafeBench offers the first comprehensive benchmark for evaluating medication safety in LLMs, advancing safer and more trustworthy AI-driven clinical decision support.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04328
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
Zhao, Jiahao
Xu, Luxin
Tan, Minghuan
Zhang, Lichao
Argha, Ahmadreza
Alinejad-Rokny, Hamid
Yang, Min
Artificial Intelligence
Numerous medical systems powered by Large Language Models (LLMs) have achieved remarkable progress in diverse healthcare tasks. However, research on their medication safety remains limited due to the lack of real world datasets, constrained by privacy and accessibility issues. Moreover, evaluation of LLMs in realistic clinical consultation settings, particularly regarding medication safety, is still underexplored. To address these gaps, we propose a framework that simulates and evaluates clinical consultations to systematically assess the medication safety capabilities of LLMs. Within this framework, we generate inquiry diagnosis dialogues with embedded medication risks and construct a dedicated medication safety database, RxRisk DB, containing 6,725 contraindications, 28,781 drug interactions, and 14,906 indication-drug pairs. A two-stage filtering strategy ensures clinical realism and professional quality, resulting in the benchmark RxSafeBench with 2,443 high-quality consultation scenarios. We evaluate leading open-source and proprietary LLMs using structured multiple choice questions that test their ability to recommend safe medications under simulated patient contexts. Results show that current LLMs struggle to integrate contraindication and interaction knowledge, especially when risks are implied rather than explicit. Our findings highlight key challenges in ensuring medication safety in LLM-based systems and provide insights into improving reliability through better prompting and task-specific tuning. RxSafeBench offers the first comprehensive benchmark for evaluating medication safety in LLMs, advancing safer and more trustworthy AI-driven clinical decision support.
title RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
topic Artificial Intelligence
url https://arxiv.org/abs/2511.04328