CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Sijia, Li, Xiaomin, Zhang, Mengxue, Jiang, Eric Hanchen, Zeng, Qingcheng, Yu, Chen-Hsiang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
par: Li, Xiaomin, et autres
Publié: (2025)
par: Li, Xiaomin, et autres
Publié: (2025)
CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
par: Xia, Peng, et autres
Publié: (2024)
par: Xia, Peng, et autres
Publié: (2024)
CSSBench: Evaluating the Safety of Lightweight LLMs against Chinese-Specific Adversarial Patterns
par: Zhou, Zhenhong, et autres
Publié: (2026)
par: Zhou, Zhenhong, et autres
Publié: (2026)
Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
par: Chen, Hongyu, et autres
Publié: (2025)
par: Chen, Hongyu, et autres
Publié: (2025)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
par: Liu, Songyang, et autres
Publié: (2025)
par: Liu, Songyang, et autres
Publié: (2025)
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
par: Wang, Shirui, et autres
Publié: (2025)
par: Wang, Shirui, et autres
Publié: (2025)
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
par: Huang, Shulin, et autres
Publié: (2025)
par: Huang, Shulin, et autres
Publié: (2025)
Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
par: Cui, Sijia, et autres
Publié: (2025)
par: Cui, Sijia, et autres
Publié: (2025)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
par: Jiang, Yilin, et autres
Publié: (2025)
par: Jiang, Yilin, et autres
Publié: (2025)
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
par: Liu, Jiayu, et autres
Publié: (2026)
par: Liu, Jiayu, et autres
Publié: (2026)
KG-Rank: Enhancing Large Language Models for Medical QA with Knowledge Graphs and Ranking Techniques
par: Yang, Rui, et autres
Publié: (2024)
par: Yang, Rui, et autres
Publié: (2024)
GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
par: Li, Qintong, et autres
Publié: (2024)
par: Li, Qintong, et autres
Publié: (2024)
CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models
par: Jia, Rui, et autres
Publié: (2026)
par: Jia, Rui, et autres
Publié: (2026)
CMB: A Comprehensive Medical Benchmark in Chinese
par: Wang, Xidong, et autres
Publié: (2023)
par: Wang, Xidong, et autres
Publié: (2023)
Large Language Models on Wikipedia-Style Survey Generation: an Evaluation in NLP Concepts
par: Gao, Fan, et autres
Publié: (2023)
par: Gao, Fan, et autres
Publié: (2023)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
par: Obadinma, Stephen, et autres
Publié: (2025)
par: Obadinma, Stephen, et autres
Publié: (2025)
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
par: Li, Xiaomin, et autres
Publié: (2025)
par: Li, Xiaomin, et autres
Publié: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
par: Hou, Yutao, et autres
Publié: (2026)
par: Hou, Yutao, et autres
Publié: (2026)
Thinking Out Loud: Do Reasoning Models Know When They're Right?
par: Zeng, Qingcheng, et autres
Publié: (2025)
par: Zeng, Qingcheng, et autres
Publié: (2025)
Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD
par: Tan, Bryan Chen Zhengyu, et autres
Publié: (2025)
par: Tan, Bryan Chen Zhengyu, et autres
Publié: (2025)
The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models
par: Yu, Kefan, et autres
Publié: (2025)
par: Yu, Kefan, et autres
Publié: (2025)
A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry
par: Huang, Yining, et autres
Publié: (2024)
par: Huang, Yining, et autres
Publié: (2024)
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
par: Zhou, Yuxuan, et autres
Publié: (2024)
par: Zhou, Yuxuan, et autres
Publié: (2024)
MedCare: Advancing Medical LLMs through Decoupling Clinical Alignment and Knowledge Aggregation
par: Liao, Yusheng, et autres
Publié: (2024)
par: Liao, Yusheng, et autres
Publié: (2024)
Data-adaptive Safety Rules for Training Reward Models
par: Li, Xiaomin, et autres
Publié: (2025)
par: Li, Xiaomin, et autres
Publié: (2025)
LLMs for Doctors: Leveraging Medical LLMs to Assist Doctors, Not Replace Them
par: Xie, Wenya, et autres
Publié: (2024)
par: Xie, Wenya, et autres
Publié: (2024)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
par: Yuan, Xiaohan, et autres
Publié: (2024)
par: Yuan, Xiaohan, et autres
Publié: (2024)
Evaluating Robustness of Generative Search Engine on Adversarial Factual Questions
par: Hu, Xuming, et autres
Publié: (2024)
par: Hu, Xuming, et autres
Publié: (2024)
Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
par: Zeng, Tong, et autres
Publié: (2025)
par: Zeng, Tong, et autres
Publié: (2025)
Evaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis
par: Chen, Andong, et autres
Publié: (2025)
par: Chen, Andong, et autres
Publié: (2025)
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
par: Qiu, Haoyi, et autres
Publié: (2025)
par: Qiu, Haoyi, et autres
Publié: (2025)
Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility
par: Lam, Suet-Ying, et autres
Publié: (2025)
par: Lam, Suet-Ying, et autres
Publié: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
par: Majumdar, Ayan, et autres
Publié: (2025)
par: Majumdar, Ayan, et autres
Publié: (2025)
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement
par: Zhang, Zhexin, et autres
Publié: (2025)
par: Zhang, Zhexin, et autres
Publié: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
par: Chen, Yen-Shan, et autres
Publié: (2026)
par: Chen, Yen-Shan, et autres
Publié: (2026)
Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications
par: Chen, Zhe, et autres
Publié: (2025)
par: Chen, Zhe, et autres
Publié: (2025)
Reasoning Robustness of LLMs to Adversarial Typographical Errors
par: Gan, Esther, et autres
Publié: (2024)
par: Gan, Esther, et autres
Publié: (2024)
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
par: Li, Daoyang, et autres
Publié: (2024)
par: Li, Daoyang, et autres
Publié: (2024)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
par: Wei, Zhang, et autres
Publié: (2025)
par: Wei, Zhang, et autres
Publié: (2025)
CFSafety: Comprehensive Fine-grained Safety Assessment for LLMs
par: Liu, Zhihao, et autres
Publié: (2024)
par: Liu, Zhihao, et autres
Publié: (2024)
Documents similaires
-
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
par: Li, Xiaomin, et autres
Publié: (2025) -
CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
par: Xia, Peng, et autres
Publié: (2024) -
CSSBench: Evaluating the Safety of Lightweight LLMs against Chinese-Specific Adversarial Patterns
par: Zhou, Zhenhong, et autres
Publié: (2026) -
Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
par: Chen, Hongyu, et autres
Publié: (2025) -
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
par: Liu, Songyang, et autres
Publié: (2025)