Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 2030
Fuente:
arXiv
Saved in:
| Main Authors: | Bian, Mouxiao, Zhang, Rongzhao, Ding, Chao, Peng, Xinwei, Xu, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
by: Ding, Jinru, et al.
Published: (2025)
by: Ding, Jinru, et al.
Published: (2025)
MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models
by: Mao, Kangkun, et al.
Published: (2025)
by: Mao, Kangkun, et al.
Published: (2025)
A Novel Ophthalmic Benchmark for Evaluating Multimodal Large Language Models with Fundus Photographs and OCT Images
by: Liang, Xiaoyi, et al.
Published: (2025)
by: Liang, Xiaoyi, et al.
Published: (2025)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
by: Xu, Run, et al.
Published: (2026)
by: Xu, Run, et al.
Published: (2026)
Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies
by: Jiang, Luyi, et al.
Published: (2025)
by: Jiang, Luyi, et al.
Published: (2025)
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information
by: Harris, Joshua, et al.
Published: (2025)
by: Harris, Joshua, et al.
Published: (2025)
MiLiC-Eval: Benchmarking Multilingual LLMs for China's Minority Languages
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation
by: Wu, Xinwei, et al.
Published: (2025)
by: Wu, Xinwei, et al.
Published: (2025)
Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study
by: He, Zhichao, et al.
Published: (2025)
by: He, Zhichao, et al.
Published: (2025)
Can Large Language Models Function as Qualified Pediatricians? A Systematic Evaluation in Real-World Clinical Contexts
by: Zhu, Siyu, et al.
Published: (2025)
by: Zhu, Siyu, et al.
Published: (2025)
CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China
by: Xu, Guixian, et al.
Published: (2025)
by: Xu, Guixian, et al.
Published: (2025)
A Comparative Analysis of Ethical and Safety Gaps in LLMs using Relative Danger Coefficient
by: Tereshchenko, Yehor, et al.
Published: (2025)
by: Tereshchenko, Yehor, et al.
Published: (2025)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
by: Guo, Ruiling, et al.
Published: (2025)
by: Guo, Ruiling, et al.
Published: (2025)
Taiwan Safety Benchmark and Breeze Guard: Toward Trustworthy AI for Taiwanese Mandarin
by: Hsu, Po-Chun, et al.
Published: (2026)
by: Hsu, Po-Chun, et al.
Published: (2026)
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
by: Zhu, Xiangyang, et al.
Published: (2025)
by: Zhu, Xiangyang, et al.
Published: (2025)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
by: Yuan, Tongxin, et al.
Published: (2024)
by: Yuan, Tongxin, et al.
Published: (2024)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
by: Liu, Chuang, et al.
Published: (2024)
by: Liu, Chuang, et al.
Published: (2024)
Healthcare Copilot: Eliciting the Power of General LLMs for Medical Consultation
by: Ren, Zhiyao, et al.
Published: (2024)
by: Ren, Zhiyao, et al.
Published: (2024)
Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer
by: Yang, Haoyan, et al.
Published: (2024)
by: Yang, Haoyan, et al.
Published: (2024)
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
by: Wang, Shirui, et al.
Published: (2025)
by: Wang, Shirui, et al.
Published: (2025)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
by: Hamna, Hamna, et al.
Published: (2025)
by: Hamna, Hamna, et al.
Published: (2025)
Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
by: Ding, Peng, et al.
Published: (2025)
by: Ding, Peng, et al.
Published: (2025)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
by: Xu, Ruiling, et al.
Published: (2025)
by: Xu, Ruiling, et al.
Published: (2025)
Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs
by: Hayakawa, Akio, et al.
Published: (2025)
by: Hayakawa, Akio, et al.
Published: (2025)
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
by: Bian, Haonan, et al.
Published: (2026)
by: Bian, Haonan, et al.
Published: (2026)
MC$^2$: Towards Transparent and Culturally-Aware NLP for Minority Languages in China
by: Zhang, Chen, et al.
Published: (2023)
by: Zhang, Chen, et al.
Published: (2023)
Locking Down the Finetuned LLMs Safety
by: Zhu, Minjun, et al.
Published: (2024)
by: Zhu, Minjun, et al.
Published: (2024)
SlangDIT: Benchmarking LLMs in Interpretative Slang Translation
by: Liang, Yunlong, et al.
Published: (2025)
by: Liang, Yunlong, et al.
Published: (2025)
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
by: Ding, Peng, et al.
Published: (2025)
by: Ding, Peng, et al.
Published: (2025)
ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents
by: Shao, Jie-Jing, et al.
Published: (2024)
by: Shao, Jie-Jing, et al.
Published: (2024)
Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
by: Hazra, Rima, et al.
Published: (2026)
by: Hazra, Rima, et al.
Published: (2026)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
by: Duan, Shitong, et al.
Published: (2023)
by: Duan, Shitong, et al.
Published: (2023)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
by: Rios-Sialer, Ian
Published: (2026)
by: Rios-Sialer, Ian
Published: (2026)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
by: Li, Guanlin, et al.
Published: (2025)
by: Li, Guanlin, et al.
Published: (2025)
Copyright Detective: A Forensic System to Evidence LLMs Flickering Copyright Leakage Risks
by: Zhang, Guangwei, et al.
Published: (2026)
by: Zhang, Guangwei, et al.
Published: (2026)
Similar Items
-
MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
by: Ding, Jinru, et al.
Published: (2025) -
MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models
by: Mao, Kangkun, et al.
Published: (2025) -
A Novel Ophthalmic Benchmark for Evaluating Multimodal Large Language Models with Fundus Photographs and OCT Images
by: Liang, Xiaoyi, et al.
Published: (2025) -
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
by: Xu, Run, et al.
Published: (2026) -
Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies
by: Jiang, Luyi, et al.
Published: (2025)