MedEthicEval: Evaluating Large Language Models Based on Chinese Medical Ethics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Haoan, Shi, Jiacheng, Xu, Hanhui, Zhu, Kenny Q., Wu, Mengyue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915181501612032
author Jin, Haoan
Shi, Jiacheng
Xu, Hanhui
Zhu, Kenny Q.
Wu, Mengyue
author_facet Jin, Haoan
Shi, Jiacheng
Xu, Hanhui
Zhu, Kenny Q.
Wu, Mengyue
contents Large language models (LLMs) demonstrate significant potential in advancing medical applications, yet their capabilities in addressing medical ethics challenges remain underexplored. This paper introduces MedEthicEval, a novel benchmark designed to systematically evaluate LLMs in the domain of medical ethics. Our framework encompasses two key components: knowledge, assessing the models' grasp of medical ethics principles, and application, focusing on their ability to apply these principles across diverse scenarios. To support this benchmark, we consulted with medical ethics researchers and developed three datasets addressing distinct ethical challenges: blatant violations of medical ethics, priority dilemmas with clear inclinations, and equilibrium dilemmas without obvious resolutions. MedEthicEval serves as a critical tool for understanding LLMs' ethical reasoning in healthcare, paving the way for their responsible and effective use in medical contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedEthicEval: Evaluating Large Language Models Based on Chinese Medical Ethics
Jin, Haoan
Shi, Jiacheng
Xu, Hanhui
Zhu, Kenny Q.
Wu, Mengyue
Computation and Language
Large language models (LLMs) demonstrate significant potential in advancing medical applications, yet their capabilities in addressing medical ethics challenges remain underexplored. This paper introduces MedEthicEval, a novel benchmark designed to systematically evaluate LLMs in the domain of medical ethics. Our framework encompasses two key components: knowledge, assessing the models' grasp of medical ethics principles, and application, focusing on their ability to apply these principles across diverse scenarios. To support this benchmark, we consulted with medical ethics researchers and developed three datasets addressing distinct ethical challenges: blatant violations of medical ethics, priority dilemmas with clear inclinations, and equilibrium dilemmas without obvious resolutions. MedEthicEval serves as a critical tool for understanding LLMs' ethical reasoning in healthcare, paving the way for their responsible and effective use in medical contexts.
title MedEthicEval: Evaluating Large Language Models Based on Chinese Medical Ethics
topic Computation and Language
url https://arxiv.org/abs/2503.02374