Assessing LLMs' Performance: Insights from the Chinese Pharmacist Exam
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xinran, Zhu, Boran, Zhou, Shujuan, Long, Ziwen, Zhou, Dehua, Zhang, Shu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CJEval: A Benchmark for Assessing Large Language Models Using Chinese Junior High School Exam Data
by: Zhang, Qian-Wen, et al.
Published: (2024)
by: Zhang, Qian-Wen, et al.
Published: (2024)
DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
by: Wan, Haiyuan, et al.
Published: (2025)
by: Wan, Haiyuan, et al.
Published: (2025)
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
by: Zhu, Yunqi, et al.
Published: (2024)
by: Zhu, Yunqi, et al.
Published: (2024)
Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities
by: Zhang, Lili, et al.
Published: (2025)
by: Zhang, Lili, et al.
Published: (2025)
Transfer Learning for Bayesian Optimization on Heterogeneous Search Spaces
by: Fan, Zhou, et al.
Published: (2023)
by: Fan, Zhou, et al.
Published: (2023)
Flames: Benchmarking Value Alignment of LLMs in Chinese
by: Huang, Kexin, et al.
Published: (2023)
by: Huang, Kexin, et al.
Published: (2023)
DROJ: A Prompt-Driven Attack against Large Language Models
by: Hu, Leyang, et al.
Published: (2024)
by: Hu, Leyang, et al.
Published: (2024)
Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study
by: Isley, Calvin, et al.
Published: (2025)
by: Isley, Calvin, et al.
Published: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025)
by: Liu, Chaoqun, et al.
Published: (2025)
KG-ASG: Collision-Knowledge-Guided Closed-Loop Adversarial Scenario Generation With Primary-Support Attribution
by: Wang, Cheng, et al.
Published: (2026)
by: Wang, Cheng, et al.
Published: (2026)
FERA: A Pose-Based Framework for Rule-Grounded Multimedia Decision Support with a Foil Fencing Case Study
by: Chen, Ziwen, et al.
Published: (2025)
by: Chen, Ziwen, et al.
Published: (2025)
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
by: Wu, Chengwei, et al.
Published: (2025)
by: Wu, Chengwei, et al.
Published: (2025)
The Ever-Evolving Science Exam
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
Enhancing few-shot time series forecasting with LLM-guided diffusion
by: Shi, Haonan, et al.
Published: (2026)
by: Shi, Haonan, et al.
Published: (2026)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
by: Li, Xinran, et al.
Published: (2025)
by: Li, Xinran, et al.
Published: (2025)
Machine-Assisted Grading of Nationwide School-Leaving Essay Exams with LLMs and Statistical NLP
by: Karjus, Andres, et al.
Published: (2026)
by: Karjus, Andres, et al.
Published: (2026)
DiffuSent: Towards a Unified Diffusion Framework for Aspect-Based Sentiment Analysis
by: Long, Shu, et al.
Published: (2026)
by: Long, Shu, et al.
Published: (2026)
InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents
by: Zhu, Zhenghao, et al.
Published: (2025)
by: Zhu, Zhenghao, et al.
Published: (2025)
Fùxì: A Benchmark for Evaluating Language Models on Ancient Chinese Text Understanding and Generation
by: Zhao, Shangqing, et al.
Published: (2025)
by: Zhao, Shangqing, et al.
Published: (2025)
SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
by: Chai, Jingyi, et al.
Published: (2025)
by: Chai, Jingyi, et al.
Published: (2025)
Humanity's Last Exam
by: Phan, Long, et al.
Published: (2025)
by: Phan, Long, et al.
Published: (2025)
RankAdaptor: Hierarchical Rank Allocation for Efficient Fine-Tuning Pruned LLMs via Performance Model
by: Zhou, Changhai, et al.
Published: (2024)
by: Zhou, Changhai, et al.
Published: (2024)
Overview of AI Grading of Physics Olympiad Exams
by: McGinness, Lachlan
Published: (2025)
by: McGinness, Lachlan
Published: (2025)
DialogueReason: Rule-Based RL Sparks Dialogue Reasoning in LLMs
by: Shu, Yubo, et al.
Published: (2025)
by: Shu, Yubo, et al.
Published: (2025)
ChineseErrorCorrector3-4B: State-of-the-Art Chinese Spelling and Grammar Corrector
by: Tian, Wei, et al.
Published: (2025)
by: Tian, Wei, et al.
Published: (2025)
Query Disambiguation via Answer-Free Context: Doubling Performance on Humanity's Last Exam
by: Majurski, Michael, et al.
Published: (2026)
by: Majurski, Michael, et al.
Published: (2026)
Human-AI Collaborative Game Testing with Vision Language Models
by: Zhang, Boran, et al.
Published: (2025)
by: Zhang, Boran, et al.
Published: (2025)
What Breaks Knowledge Graph based RAG? Benchmarking and Empirical Insights into Reasoning under Incomplete Knowledge
by: Zhou, Dongzhuoran, et al.
Published: (2025)
by: Zhou, Dongzhuoran, et al.
Published: (2025)
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
by: Zhao, Jinghua, et al.
Published: (2025)
by: Zhao, Jinghua, et al.
Published: (2025)
Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
by: Zhou, Yuhao, et al.
Published: (2025)
by: Zhou, Yuhao, et al.
Published: (2025)
FinLLMs: A Framework for Financial Reasoning Dataset Generation with Large Language Models
by: Yuan, Ziqiang, et al.
Published: (2024)
by: Yuan, Ziqiang, et al.
Published: (2024)
medicX-KG: A Knowledge Graph for Pharmacists' Drug Information Needs
by: Farrugia, Lizzy, et al.
Published: (2025)
by: Farrugia, Lizzy, et al.
Published: (2025)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
by: Zhou, Chengliang, et al.
Published: (2025)
by: Zhou, Chengliang, et al.
Published: (2025)
SMARTAPS: Tool-augmented LLMs for Operations Management
by: Yu, Timothy Tin Long, et al.
Published: (2025)
by: Yu, Timothy Tin Long, et al.
Published: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
Iterative Semantic Reasoning from Individual to Group Interests for Generative Recommendation with LLMs
by: Zhu, Xiaofei, et al.
Published: (2026)
by: Zhu, Xiaofei, et al.
Published: (2026)
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
CTourLLM: Enhancing LLMs with Chinese Tourism Knowledge
by: Wei, Qikai, et al.
Published: (2024)
by: Wei, Qikai, et al.
Published: (2024)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
by: Wang, Shanshan, et al.
Published: (2025)
by: Wang, Shanshan, et al.
Published: (2025)
Similar Items
-
CJEval: A Benchmark for Assessing Large Language Models Using Chinese Junior High School Exam Data
by: Zhang, Qian-Wen, et al.
Published: (2024) -
DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
by: Wan, Haiyuan, et al.
Published: (2025) -
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
by: Zhu, Yunqi, et al.
Published: (2024) -
Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities
by: Zhang, Lili, et al.
Published: (2025) -
Transfer Learning for Bayesian Optimization on Heterogeneous Search Spaces
by: Fan, Zhou, et al.
Published: (2023)