MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Corbeil, Jean-Philippe, Kim, Minseon, Griot, Maxime, Agarwal, Sheela, Sordoni, Alessandro, Beaulieu, Francois, Vozila, Paul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911361661927424
author Corbeil, Jean-Philippe
Kim, Minseon
Griot, Maxime
Agarwal, Sheela
Sordoni, Alessandro
Beaulieu, Francois
Vozila, Paul
author_facet Corbeil, Jean-Philippe
Kim, Minseon
Griot, Maxime
Agarwal, Sheela
Sordoni, Alessandro
Beaulieu, Francois
Vozila, Paul
contents As the performance of large language models (LLMs) continues to advance, their adoption in the medical domain is increasing. However, most existing risk evaluations largely focused on general safety benchmarks. In the medical applications, LLMs may be used by a wide range of users, ranging from general users and patients to clinicians, with diverse levels of expertise and the model's outputs can have a direct impact on human health which raises serious safety concerns. In this paper, we introduce MedRiskEval, a medical risk evaluation benchmark tailored to the medical domain. To fill the gap in previous benchmarks that only focused on the clinician perspective, we introduce a new patient-oriented dataset called PatientSafetyBench containing 466 samples across 5 critical risk categories. Leveraging our new benchmark alongside existing datasets, we evaluate a variety of open- and closed-source LLMs. To the best of our knowledge, this work establishes an initial foundation for safer deployment of LLMs in healthcare.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07248
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings
Corbeil, Jean-Philippe
Kim, Minseon
Griot, Maxime
Agarwal, Sheela
Sordoni, Alessandro
Beaulieu, Francois
Vozila, Paul
Computation and Language
As the performance of large language models (LLMs) continues to advance, their adoption in the medical domain is increasing. However, most existing risk evaluations largely focused on general safety benchmarks. In the medical applications, LLMs may be used by a wide range of users, ranging from general users and patients to clinicians, with diverse levels of expertise and the model's outputs can have a direct impact on human health which raises serious safety concerns. In this paper, we introduce MedRiskEval, a medical risk evaluation benchmark tailored to the medical domain. To fill the gap in previous benchmarks that only focused on the clinician perspective, we introduce a new patient-oriented dataset called PatientSafetyBench containing 466 samples across 5 critical risk categories. Leveraging our new benchmark alongside existing datasets, we evaluate a variety of open- and closed-source LLMs. To the best of our knowledge, this work establishes an initial foundation for safer deployment of LLMs in healthcare.
title MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings
topic Computation and Language
url https://arxiv.org/abs/2507.07248