LLM Sensitivity Evaluation Framework for Clinical Diagnosis
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Chenwei, Fu, Xiangling, Xiong, Yuxuan, Wang, Tianyi, Hui, Siu Cheung, Wu, Ji, Liu, Xien |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
Data Augmentation Techniques for Chinese Disease Name Normalization
by: Cui, Wenqian, et al.
Published: (2025)
by: Cui, Wenqian, et al.
Published: (2025)
Simple Data Augmentation Techniques for Chinese Disease Normalization
by: Cui, Wenqian, et al.
Published: (2023)
by: Cui, Wenqian, et al.
Published: (2023)
Reliable and diverse evaluation of LLM medical knowledge mastery
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
A Structure-aware Generative Model for Biomedical Event Extraction
by: Yuan, Haohan, et al.
Published: (2024)
by: Yuan, Haohan, et al.
Published: (2024)
Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
MKE-Coder: Multi-Axial Knowledge with Evidence Verification in ICD Coding for Chinese EMRs
by: You, Xinxin, et al.
Published: (2025)
by: You, Xinxin, et al.
Published: (2025)
TinyLLaVA: A Framework of Small-scale Large Multimodal Models
by: Zhou, Baichuan, et al.
Published: (2024)
by: Zhou, Baichuan, et al.
Published: (2024)
Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs
by: Zheng, Huan, et al.
Published: (2026)
by: Zheng, Huan, et al.
Published: (2026)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Unintended Negative Impacts of Promotional Language in Patent Evaluation
by: Zhao, Bingkun, et al.
Published: (2026)
by: Zhao, Bingkun, et al.
Published: (2026)
Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging
by: Wu, Han, et al.
Published: (2025)
by: Wu, Han, et al.
Published: (2025)
Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization
by: Liu, Weixin, et al.
Published: (2026)
by: Liu, Weixin, et al.
Published: (2026)
medIKAL: Integrating Knowledge Graphs as Assistants of LLMs for Enhanced Clinical Diagnosis on EMRs
by: Jia, Mingyi, et al.
Published: (2024)
by: Jia, Mingyi, et al.
Published: (2024)
Rethinking LLM Ensembling from the Perspective of Mixture Models
by: Fu, Jiale, et al.
Published: (2026)
by: Fu, Jiale, et al.
Published: (2026)
DongYuan: An LLM-Based Framework for Integrative Chinese and Western Medicine Spleen-Stomach Disorders Diagnosis
by: Li, Hua, et al.
Published: (2026)
by: Li, Hua, et al.
Published: (2026)
DistillNote: Toward a Functional Evaluation Framework of LLM-Generated Clinical Note Summaries
by: Boll, Heloisa Oss, et al.
Published: (2025)
by: Boll, Heloisa Oss, et al.
Published: (2025)
A Reality check of the benefits of LLM in business
by: Cheung, Ming
Published: (2024)
by: Cheung, Ming
Published: (2024)
Human-AI Co-reasoning for Clinical Diagnosis with Evidence-Integrated Language Agent
by: Huang, Zhongzhen, et al.
Published: (2026)
by: Huang, Zhongzhen, et al.
Published: (2026)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
by: Wu, JiaRu, et al.
Published: (2025)
by: Wu, JiaRu, et al.
Published: (2025)
Merging Beyond: Streaming LLM Updates via Activation-Guided Rotations
by: Yao, Yuxuan, et al.
Published: (2026)
by: Yao, Yuxuan, et al.
Published: (2026)
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
by: Li, Yifei, et al.
Published: (2026)
by: Li, Yifei, et al.
Published: (2026)
InfiFusion: A Unified Framework for Enhanced Cross-Model Reasoning via LLM Fusion
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
Chain-of-Thought Reasoning with Large Language Models for Clinical Alzheimer's Disease Assessment and Diagnosis
by: Zhang, Tongze, et al.
Published: (2026)
by: Zhang, Tongze, et al.
Published: (2026)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Neutralizing Bias in LLM Reasoning using Entailment Graphs
by: Cheng, Liang, et al.
Published: (2025)
by: Cheng, Liang, et al.
Published: (2025)
A Specialized Large Language Model for Clinical Reasoning and Diagnosis in Rare Diseases
by: Yang, Tao, et al.
Published: (2025)
by: Yang, Tao, et al.
Published: (2025)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
by: Shin, Hyungyu, et al.
Published: (2025)
by: Shin, Hyungyu, et al.
Published: (2025)
Automating Intervention Discovery from Scientific Literature: A Progressive Ontology Prompting and Dual-LLM Framework
by: Hu, Yuting, et al.
Published: (2024)
by: Hu, Yuting, et al.
Published: (2024)
Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis
by: Wang, Haochun, et al.
Published: (2024)
by: Wang, Haochun, et al.
Published: (2024)
PsycoLLM: Enhancing LLM for Psychological Understanding and Evaluation
by: Hu, Jinpeng, et al.
Published: (2024)
by: Hu, Jinpeng, et al.
Published: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
by: Zhou, Yilun, et al.
Published: (2025)
by: Zhou, Yilun, et al.
Published: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Analyzing Diversity in Healthcare LLM Research: A Scientometric Perspective
by: Restrepo, David, et al.
Published: (2024)
by: Restrepo, David, et al.
Published: (2024)
Knowledge Reasoning of Large Language Models Integrating Graph-Structured Information for Pest and Disease Control in Tobacco
by: Li, Siyu, et al.
Published: (2025)
by: Li, Siyu, et al.
Published: (2025)
Text Classification Based on Knowledge Graphs and Improved Attention Mechanism
by: Li, Siyu, et al.
Published: (2024)
by: Li, Siyu, et al.
Published: (2024)
Craw4LLM: Efficient Web Crawling for LLM Pretraining
by: Yu, Shi, et al.
Published: (2025)
by: Yu, Shi, et al.
Published: (2025)
Similar Items
-
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
by: Zhou, Yuxuan, et al.
Published: (2025) -
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
by: Zhou, Yuxuan, et al.
Published: (2024) -
Data Augmentation Techniques for Chinese Disease Name Normalization
by: Cui, Wenqian, et al.
Published: (2025) -
Simple Data Augmentation Techniques for Chinese Disease Normalization
by: Cui, Wenqian, et al.
Published: (2023) -
Reliable and diverse evaluation of LLM medical knowledge mastery
by: Zhou, Yuxuan, et al.
Published: (2024)