Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yuxuan, Liu, Xien, Yan, Chenwei, Ning, Chen, Zhang, Xiao, Li, Boxun, Fu, Xiangling, Wang, Shijin, Hu, Guoping, Wang, Yu, Wu, Ji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
LLM Sensitivity Evaluation Framework for Clinical Diagnosis
by: Yan, Chenwei, et al.
Published: (2025)
by: Yan, Chenwei, et al.
Published: (2025)
Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
by: Cui, Yiming, et al.
Published: (2025)
by: Cui, Yiming, et al.
Published: (2025)
Reliable and diverse evaluation of LLM medical knowledge mastery
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
Data Augmentation Techniques for Chinese Disease Name Normalization
by: Cui, Wenqian, et al.
Published: (2025)
by: Cui, Wenqian, et al.
Published: (2025)
Simple Data Augmentation Techniques for Chinese Disease Normalization
by: Cui, Wenqian, et al.
Published: (2023)
by: Cui, Wenqian, et al.
Published: (2023)
ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
by: Wang, Xingqi, et al.
Published: (2025)
by: Wang, Xingqi, et al.
Published: (2025)
From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions
by: Qu, Changle, et al.
Published: (2024)
by: Qu, Changle, et al.
Published: (2024)
Learning to Solve Geometry Problems via Simulating Human Dual-Reasoning Process
by: Xiao, Tong, et al.
Published: (2024)
by: Xiao, Tong, et al.
Published: (2024)
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
by: Kale, Sahil
Published: (2025)
by: Kale, Sahil
Published: (2025)
MedM-VL: What Makes a Good Medical LVLM?
by: Shi, Yiming, et al.
Published: (2025)
by: Shi, Yiming, et al.
Published: (2025)
KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing
by: Li, Zhifei, et al.
Published: (2025)
by: Li, Zhifei, et al.
Published: (2025)
MKE-Coder: Multi-Axial Knowledge with Evidence Verification in ICD Coding for Chinese EMRs
by: You, Xinxin, et al.
Published: (2025)
by: You, Xinxin, et al.
Published: (2025)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
by: Wang, Zihu, et al.
Published: (2025)
by: Wang, Zihu, et al.
Published: (2025)
ADO-LLM: Analog Design Bayesian Optimization with In-Context Learning of Large Language Models
by: Yin, Yuxuan, et al.
Published: (2024)
by: Yin, Yuxuan, et al.
Published: (2024)
From Beginner to Expert: Modeling Medical Knowledge into General LLMs
by: Li, Qiang, et al.
Published: (2023)
by: Li, Qiang, et al.
Published: (2023)
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
by: Wang, Xiaoxuan, et al.
Published: (2023)
by: Wang, Xiaoxuan, et al.
Published: (2023)
From Conversation to Automation: Leveraging LLMs for Problem-Solving Therapy Analysis
by: Aghakhani, Elham, et al.
Published: (2025)
by: Aghakhani, Elham, et al.
Published: (2025)
Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
by: Li, Alan, et al.
Published: (2025)
by: Li, Alan, et al.
Published: (2025)
Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
Human Cognition Inspired RAG with Knowledge Graph for Complex Problem Solving
by: Cheng, Yao, et al.
Published: (2025)
by: Cheng, Yao, et al.
Published: (2025)
Logic of (Common or Distributed) Knowledge
by: Shi, Chenwei
Published: (2025)
by: Shi, Chenwei
Published: (2025)
Can LLMs Solve longer Math Word Problems Better?
by: Xu, Xin, et al.
Published: (2024)
by: Xu, Xin, et al.
Published: (2024)
Bridging the Knowledge-Action Gap by Evaluating LLMs in Dynamic Dental Clinical Scenarios
by: Ma, Hongyang, et al.
Published: (2026)
by: Ma, Hongyang, et al.
Published: (2026)
From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning
by: Dinucu-Jianu, David, et al.
Published: (2025)
by: Dinucu-Jianu, David, et al.
Published: (2025)
From Failure to Mastery: Generating Hard Samples for Tool-use Agents
by: Hao, Bingguang, et al.
Published: (2026)
by: Hao, Bingguang, et al.
Published: (2026)
Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
by: Xu, Ning, et al.
Published: (2025)
by: Xu, Ning, et al.
Published: (2025)
Enhancing Paediatric Deterioration Assessment Across Diverse Skin Tones: Insights and Future Directions
by: Tianji Zhou, et al.
Published: (2025)
by: Tianji Zhou, et al.
Published: (2025)
Evaluating Developmental Cognition Capabilities of LLMs
by: Xiao, Xiao, et al.
Published: (2026)
by: Xiao, Xiao, et al.
Published: (2026)
Hán Dān Xué Bù (Mimicry) or Qīng Chū Yú Lán (Mastery)? A Cognitive Perspective on Reasoning Distillation in Large Language Models
by: Hu, Yueqing, et al.
Published: (2026)
by: Hu, Yueqing, et al.
Published: (2026)
A Generalized Additive Partial-Mastery Cognitive Diagnosis Model
by: Cárdenas-Hurtado, Camilo, et al.
Published: (2025)
by: Cárdenas-Hurtado, Camilo, et al.
Published: (2025)
CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling
by: Wang, Bichen, et al.
Published: (2025)
by: Wang, Bichen, et al.
Published: (2025)
A Concept for Autonomous Problem-Solving in Intralogistics Scenarios
by: Sigel, Johannes, et al.
Published: (2025)
by: Sigel, Johannes, et al.
Published: (2025)
Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator
by: Liu, Chengyuan, et al.
Published: (2024)
by: Liu, Chengyuan, et al.
Published: (2024)
From Correction to Mastery: Reinforced Distillation of Large Language Model Agents
by: Lyu, Yuanjie, et al.
Published: (2025)
by: Lyu, Yuanjie, et al.
Published: (2025)
CreDes: Causal Reasoning Enhancement and Dual-End Searching for Solving Long-Range Reasoning Problems using LLMs
by: Wang, Kangsheng, et al.
Published: (2024)
by: Wang, Kangsheng, et al.
Published: (2024)
CE-GOCD: Central Entity-Guided Graph Optimization for Community Detection to Augment LLM Scientific Question Answering
by: Lan, Jiayin, et al.
Published: (2026)
by: Lan, Jiayin, et al.
Published: (2026)
MultiMedEdit: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQA
by: Wen, Shengtao, et al.
Published: (2025)
by: Wen, Shengtao, et al.
Published: (2025)
GELD: A Unified Neural Model for Efficiently Solving Traveling Salesman Problems Across Different Scales
by: Xiao, Yubin, et al.
Published: (2025)
by: Xiao, Yubin, et al.
Published: (2025)
Similar Items
-
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
by: Zhou, Yuxuan, et al.
Published: (2024) -
LLM Sensitivity Evaluation Framework for Clinical Diagnosis
by: Yan, Chenwei, et al.
Published: (2025) -
Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
by: Zhou, Yuxuan, et al.
Published: (2025) -
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
by: Cui, Yiming, et al.
Published: (2025) -
Reliable and diverse evaluation of LLM medical knowledge mastery
by: Zhou, Yuxuan, et al.
Published: (2024)