Reliable and diverse evaluation of LLM medical knowledge mastery
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yuxuan, Liu, Xien, Ning, Chen, Zhang, Xiao, Wu, Ji |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
MKE-Coder: Multi-Axial Knowledge with Evidence Verification in ICD Coding for Chinese EMRs
by: You, Xinxin, et al.
Published: (2025)
by: You, Xinxin, et al.
Published: (2025)
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
by: Xiao, Yilin, et al.
Published: (2025)
by: Xiao, Yilin, et al.
Published: (2025)
Data Augmentation Techniques for Chinese Disease Name Normalization
by: Cui, Wenqian, et al.
Published: (2025)
by: Cui, Wenqian, et al.
Published: (2025)
Simple Data Augmentation Techniques for Chinese Disease Normalization
by: Cui, Wenqian, et al.
Published: (2023)
by: Cui, Wenqian, et al.
Published: (2023)
A unified foundational framework for knowledge injection and evaluation of Large Language Models in Combustion Science
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
Evidence of conceptual mastery in the application of rules by Large Language Models
by: Nunes, José Luiz, et al.
Published: (2025)
by: Nunes, José Luiz, et al.
Published: (2025)
How well do LLMs cite relevant medical references? An evaluation framework and analyses
by: Wu, Kevin, et al.
Published: (2024)
by: Wu, Kevin, et al.
Published: (2024)
Polish-English medical knowledge transfer: A new benchmark and results
by: Grzybowski, Łukasz, et al.
Published: (2024)
by: Grzybowski, Łukasz, et al.
Published: (2024)
ADMEDTAGGER: an annotation framework for distillation of expert knowledge for the Polish medical language
by: Górski, Franciszek, et al.
Published: (2025)
by: Górski, Franciszek, et al.
Published: (2025)
Reasoning based on symbolic and parametric knowledge bases: a survey
by: Xu, Mayi, et al.
Published: (2025)
by: Xu, Mayi, et al.
Published: (2025)
Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
How new data permeates LLM knowledge and how to dilute it
by: Sun, Chen, et al.
Published: (2025)
by: Sun, Chen, et al.
Published: (2025)
keqing: knowledge-based question answering is a nature chain-of-thought mentor of LLM
by: Wang, Chaojie, et al.
Published: (2023)
by: Wang, Chaojie, et al.
Published: (2023)
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
by: Chen, Zichen, et al.
Published: (2023)
by: Chen, Zichen, et al.
Published: (2023)
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
by: Yuan, Peiwen, et al.
Published: (2025)
by: Yuan, Peiwen, et al.
Published: (2025)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
by: Wang, Guanghui, et al.
Published: (2025)
by: Wang, Guanghui, et al.
Published: (2025)
CVE-LLM : Automatic vulnerability evaluation in medical device industry using large language models
by: Ghosh, Rikhiya, et al.
Published: (2024)
by: Ghosh, Rikhiya, et al.
Published: (2024)
CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
by: Zhu, Wenqiao, et al.
Published: (2025)
by: Zhu, Wenqiao, et al.
Published: (2025)
LLM-based Prompt Ensemble for Reliable Medical Entity Recognition from EHRs
by: Islam, K M Sajjadul, et al.
Published: (2025)
by: Islam, K M Sajjadul, et al.
Published: (2025)
Planning in the LLM Era: Building for Reliability and Efficiency
by: Katz, Michael, et al.
Published: (2026)
by: Katz, Michael, et al.
Published: (2026)
An LLM Maturity Model for Reliable and Transparent Text-to-Query
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
by: Gonzalez, Miguel Angel Alvarado, et al.
Published: (2025)
by: Gonzalez, Miguel Angel Alvarado, et al.
Published: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
by: Bologna, Federica, et al.
Published: (2026)
by: Bologna, Federica, et al.
Published: (2026)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
by: Sands, Brendan, et al.
Published: (2025)
by: Sands, Brendan, et al.
Published: (2025)
Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT
by: Tao, Zhen, et al.
Published: (2024)
by: Tao, Zhen, et al.
Published: (2024)
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
by: Wang, Shaojie, et al.
Published: (2026)
by: Wang, Shaojie, et al.
Published: (2026)
Evaluating how LLM annotations represent diverse views on contentious topics
by: Brown, Megan A., et al.
Published: (2025)
by: Brown, Megan A., et al.
Published: (2025)
Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
by: Zheng, Hang, et al.
Published: (2025)
by: Zheng, Hang, et al.
Published: (2025)
Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation
by: Wysocka, Magdalena, et al.
Published: (2023)
by: Wysocka, Magdalena, et al.
Published: (2023)
PersLLM: A Personified Training Approach for Large Language Models
by: Zeng, Zheni, et al.
Published: (2024)
by: Zeng, Zheni, et al.
Published: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
by: Krishnappa, Pushwitha, et al.
Published: (2026)
by: Krishnappa, Pushwitha, et al.
Published: (2026)
LANID: LLM-assisted New Intent Discovery
by: Fan, Lu, et al.
Published: (2025)
by: Fan, Lu, et al.
Published: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
Similar Items
-
Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
by: Zhou, Yuxuan, et al.
Published: (2025) -
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
by: Zhou, Yuxuan, et al.
Published: (2025) -
MKE-Coder: Multi-Axial Knowledge with Evidence Verification in ICD Coding for Chinese EMRs
by: You, Xinxin, et al.
Published: (2025) -
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
by: Zhou, Yuxuan, et al.
Published: (2024) -
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
by: Xiao, Yilin, et al.
Published: (2025)