Can Large Language Models Function as Qualified Pediatricians? A Systematic Evaluation in Real-World Clinical Contexts
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Siyu, Bian, Mouxiao, Xie, Yue, Tang, Yongyu, Yu, Zhikang, Li, Tianbin, Chen, Pengcheng, Han, Bing, Xu, Jie, Dong, Xiaoyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study
by: He, Zhichao, et al.
Published: (2025)
by: He, Zhichao, et al.
Published: (2025)
Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
by: Yao, Ben, et al.
Published: (2025)
by: Yao, Ben, et al.
Published: (2025)
TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine
by: Huang, Tianai, et al.
Published: (2025)
by: Huang, Tianai, et al.
Published: (2025)
Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI
by: Ding, Chao, et al.
Published: (2025)
by: Ding, Chao, et al.
Published: (2025)
A Novel Ophthalmic Benchmark for Evaluating Multimodal Large Language Models with Fundus Photographs and OCT Images
by: Liang, Xiaoyi, et al.
Published: (2025)
by: Liang, Xiaoyi, et al.
Published: (2025)
Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 2030
by: Bian, Mouxiao, et al.
Published: (2025)
by: Bian, Mouxiao, et al.
Published: (2025)
Classification of Rational Functions of Degree Three over Finite Fields
by: Hou, Xiang-dong, et al.
Published: (2026)
by: Hou, Xiang-dong, et al.
Published: (2026)
Pruning Graphs by Adversarial Robustness Evaluation to Strengthen GNN Defenses
by: Wang, Yongyu
Published: (2025)
by: Wang, Yongyu
Published: (2025)
Enabling DBSCAN for Very Large-Scale High-Dimensional Spaces
by: Wang, Yongyu
Published: (2024)
by: Wang, Yongyu
Published: (2024)
Evaluating the Usability of Qualified Electronic Signatures: Systematized Use Cases and Design Paradigms
by: Cagal, Mustafa, et al.
Published: (2024)
by: Cagal, Mustafa, et al.
Published: (2024)
Diagnostic Challenges in Pediatric Dermatologic Presentations Among Pediatricians
by: Kathy Boutis, et al.
Published: (2026)
by: Kathy Boutis, et al.
Published: (2026)
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
by: Zhang, Danyang, et al.
Published: (2023)
by: Zhang, Danyang, et al.
Published: (2023)
MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models
by: Mao, Kangkun, et al.
Published: (2025)
by: Mao, Kangkun, et al.
Published: (2025)
Can Large Language Models Detect Real-World Android Software Compliance Violations?
by: Zhang, Haoyi, et al.
Published: (2025)
by: Zhang, Haoyi, et al.
Published: (2025)
SEAL: A Framework for Systematic Evaluation of Real-World Super-Resolution
by: Zhang, Wenlong, et al.
Published: (2023)
by: Zhang, Wenlong, et al.
Published: (2023)
ECG-WM: A Physiology-Informed ECG World Model for Clinical Intervention Simulation
by: Chen, Zhikang, et al.
Published: (2026)
by: Chen, Zhikang, et al.
Published: (2026)
Can Large Language Models Understand Real-World Complex Instructions?
by: He, Qianyu, et al.
Published: (2023)
by: He, Qianyu, et al.
Published: (2023)
Mitigating the Impact of Noisy Edges on Graph-Based Algorithms via Adversarial Robustness Evaluation
by: Wang, Yongyu, et al.
Published: (2024)
by: Wang, Yongyu, et al.
Published: (2024)
Is Continual Learning Ready for Real-world Challenges?
by: Kontogianni, Theodora, et al.
Published: (2024)
by: Kontogianni, Theodora, et al.
Published: (2024)
Inflamed or Infected Molluscum Contagiosum Lesions: Pediatrician Perceptions and the Risk of Antibiotic Overuse
by: Trevor Young, et al.
Published: (2025)
by: Trevor Young, et al.
Published: (2025)
Food Allergy Diagnosis in Early Childhood: Journey Mapping Study With Parents and Pediatricians
by: Madlen Hörold, et al.
Published: (2026)
by: Madlen Hörold, et al.
Published: (2026)
EVALUATING THE PERFORMANCE OF PARALLEL COMPUTING IN HYBRID MODELS
by: Chen Yongyu
Published: (2023)
by: Chen Yongyu
Published: (2023)
Spectral structural distortion reveals redundant neurons in neural networks
by: Wang, Yongyu
Published: (2026)
by: Wang, Yongyu
Published: (2026)
Defending Collaborative Filtering Recommenders via Adversarial Robustness Based Edge Reweighting
by: Wang, Yongyu
Published: (2024)
by: Wang, Yongyu
Published: (2024)
T2T-LA: A Topology-to-Topology LLM Agent for Graph Learning with Neither Feature Access nor Task Knowledge
by: Wang, Yongyu
Published: (2025)
by: Wang, Yongyu
Published: (2025)
Adversarial-Robustness-Guided Graph Pruning
by: Wang, Yongyu
Published: (2024)
by: Wang, Yongyu
Published: (2024)
Improving Collaborative Filtering Recommendation via Graph Learning
by: Wang, Yongyu
Published: (2023)
by: Wang, Yongyu
Published: (2023)
Approximation Algorithms for Capacitated Vehicle Routing Problems: A Comprehensive Survey
by: Chen, Yongyu
Published: (2023)
by: Chen, Yongyu
Published: (2023)
Spectrally unstable nodes drive reliability failures in graph learning
by: Wang, Yongyu
Published: (2024)
by: Wang, Yongyu
Published: (2024)
Translate-and-Revise: Boosting Large Language Models for Constrained Translation
by: Huang, Pengcheng, et al.
Published: (2024)
by: Huang, Pengcheng, et al.
Published: (2024)
Evaluation of Performance Measures for Qualifying Flood Models with Satellite Observations
by: Travert, Jean-Paul, et al.
Published: (2024)
by: Travert, Jean-Paul, et al.
Published: (2024)
From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models
by: Chen, Zhikang, et al.
Published: (2026)
by: Chen, Zhikang, et al.
Published: (2026)
Evaluating LLMs Code Reasoning Under Real-World Context
by: Liu, Changshu
Published: (2026)
by: Liu, Changshu
Published: (2026)
Augmented Deep Contexts for Spatially Embedded Video Coding
by: Bian, Yifan, et al.
Published: (2025)
by: Bian, Yifan, et al.
Published: (2025)
Dual-Imbalance Continual Learning for Real-World Food Recognition
by: Zhang, Xiaoyan, et al.
Published: (2026)
by: Zhang, Xiaoyan, et al.
Published: (2026)
CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research?
by: Chen, Xiangsen, et al.
Published: (2026)
by: Chen, Xiangsen, et al.
Published: (2026)
Mechanical Exfoliation as Adjunct Therapy in Pediatric Seborrheic Dermatitis: Device Characterization and Clinical Rationale for a Pediatrician-Developed Silicone Scalp Brush
by: Valenzuela, Eduard
Published: (2026)
by: Valenzuela, Eduard
Published: (2026)
Pediatricians' practices and knowledge of metabolic dysfunction‐associated steatotic liver disease: An international survey
by: Judith W. Lubrecht, et al.
Published: (2024)
by: Judith W. Lubrecht, et al.
Published: (2024)
Barriers to Accurate Diagnosis of Infantile Atopic Dermatitis: Insights From a Survey of Pediatricians
by: Kiwako Yamamoto‐Hanada, et al.
Published: (2025)
by: Kiwako Yamamoto‐Hanada, et al.
Published: (2025)
Similar Items
-
Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study
by: He, Zhichao, et al.
Published: (2025) -
Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review
by: Yang, Yan, et al.
Published: (2025) -
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
by: Yao, Ben, et al.
Published: (2025) -
TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine
by: Huang, Tianai, et al.
Published: (2025) -
Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI
by: Ding, Chao, et al.
Published: (2025)