ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Canyu, Yu, Jian, Chen, Shan, Liu, Che, Wan, Zhongwei, Bitterman, Danielle, Wang, Fei, Shu, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV
by: Stinard, Alex
Published: (2026)
by: Stinard, Alex
Published: (2026)
Can Large Language Models Identify Authorship?
by: Huang, Baixiang, et al.
Published: (2024)
by: Huang, Baixiang, et al.
Published: (2024)
Can LLM-Generated Misinformation Be Detected?
by: Chen, Canyu, et al.
Published: (2023)
by: Chen, Canyu, et al.
Published: (2023)
Improving Clinical NLP Performance through Language Model-Generated Synthetic Clinical Data
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
Can Knowledge Editing Really Correct Hallucinations?
by: Huang, Baixiang, et al.
Published: (2024)
by: Huang, Baixiang, et al.
Published: (2024)
Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?
by: Seah, Natalie, et al.
Published: (2026)
by: Seah, Natalie, et al.
Published: (2026)
debiaSAE: Benchmarking and Mitigating Vision-Language Model Bias
by: Sasse, Kuleen, et al.
Published: (2024)
by: Sasse, Kuleen, et al.
Published: (2024)
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
by: Qi, Jirui, et al.
Published: (2025)
by: Qi, Jirui, et al.
Published: (2025)
Can Editing LLMs Inject Harm?
by: Chen, Canyu, et al.
Published: (2024)
by: Chen, Canyu, et al.
Published: (2024)
Gender Bias in Large Language Models for Healthcare: Assignment Consistency and Clinical Implications
by: Liu, Mingxuan, et al.
Published: (2025)
by: Liu, Mingxuan, et al.
Published: (2025)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
Model Attribution in LLM-Generated Disinformation: A Domain Generalization Approach with Supervised Contrastive Learning
by: Beigi, Alimohammad, et al.
Published: (2024)
by: Beigi, Alimohammad, et al.
Published: (2024)
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
by: Ye, Bingyang, et al.
Published: (2026)
by: Ye, Bingyang, et al.
Published: (2026)
RareBench: Can LLMs Serve as Rare Diseases Specialists?
by: Chen, Xuanzhong, et al.
Published: (2024)
by: Chen, Xuanzhong, et al.
Published: (2024)
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025)
by: Gallifant, Jack, et al.
Published: (2025)
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications
by: Liu, Yanchen, et al.
Published: (2023)
by: Liu, Yanchen, et al.
Published: (2023)
Jailbreak Detection in Clinical Training LLMs Using Feature-Based Predictive Models
by: Nguyen, Tri, et al.
Published: (2025)
by: Nguyen, Tri, et al.
Published: (2025)
Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
by: Chen, Shan, et al.
Published: (2023)
by: Chen, Shan, et al.
Published: (2023)
Authorship Attribution in the Era of LLMs: Problems, Methodologies, and Challenges
by: Huang, Baixiang, et al.
Published: (2024)
by: Huang, Baixiang, et al.
Published: (2024)
Can Reasoning LLMs Enhance Clinical Document Classification?
by: Mustafa, Akram, et al.
Published: (2025)
by: Mustafa, Akram, et al.
Published: (2025)
Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
by: Li, Yuchong, et al.
Published: (2025)
by: Li, Yuchong, et al.
Published: (2025)
CliniBench: A Clinical Outcome Prediction Benchmark for Generative and Encoder-Based Language Models
by: Grundmann, Paul, et al.
Published: (2025)
by: Grundmann, Paul, et al.
Published: (2025)
The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities
by: Research, MediaTek, et al.
Published: (2025)
by: Research, MediaTek, et al.
Published: (2025)
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
by: Xiao, Yunpeng, et al.
Published: (2025)
by: Xiao, Yunpeng, et al.
Published: (2025)
From Bench to Bedside: A Review of Clinical Trials in Drug Discovery and Development
by: Wang, Tianyang, et al.
Published: (2024)
by: Wang, Tianyang, et al.
Published: (2024)
Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate
by: Zhang, Yiqun, et al.
Published: (2024)
by: Zhang, Yiqun, et al.
Published: (2024)
When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?
by: Gao, Yanjun, et al.
Published: (2024)
by: Gao, Yanjun, et al.
Published: (2024)
FilBench: Can LLMs Understand and Generate Filipino?
by: Miranda, Lester James V., et al.
Published: (2025)
by: Miranda, Lester James V., et al.
Published: (2025)
ExpressivityBench: Can LLMs Communicate Implicitly?
by: Tint, Joshua, et al.
Published: (2024)
by: Tint, Joshua, et al.
Published: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
Uncertainty Quantification for Clinical Outcome Predictions with (Large) Language Models
by: Chen, Zizhang, et al.
Published: (2024)
by: Chen, Zizhang, et al.
Published: (2024)
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)
by: Zhao, Yunhan, et al.
Published: (2026)
Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
by: Gallifant, Jack, et al.
Published: (2024)
by: Gallifant, Jack, et al.
Published: (2024)
LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs
by: Lian, Da-Chen, et al.
Published: (2025)
by: Lian, Da-Chen, et al.
Published: (2025)
Leveraging LLMs for Predicting Unknown Diagnoses from Clinical Notes
by: Albassam, Dina, et al.
Published: (2025)
by: Albassam, Dina, et al.
Published: (2025)
Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability
by: Gao, Yanjun, et al.
Published: (2024)
by: Gao, Yanjun, et al.
Published: (2024)
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
by: Chen, Shan, et al.
Published: (2025)
by: Chen, Shan, et al.
Published: (2025)
From Metaphor to Mechanism: How LLMs Decode Traditional Chinese Medicine Symbolic Language for Modern Clinical Relevance
by: Tang, Jiacheng, et al.
Published: (2025)
by: Tang, Jiacheng, et al.
Published: (2025)
Similar Items
-
ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV
by: Stinard, Alex
Published: (2026) -
Can Large Language Models Identify Authorship?
by: Huang, Baixiang, et al.
Published: (2024) -
Can LLM-Generated Misinformation Be Detected?
by: Chen, Canyu, et al.
Published: (2023) -
Improving Clinical NLP Performance through Language Model-Generated Synthetic Clinical Data
by: Chen, Shan, et al.
Published: (2024) -
Can Knowledge Editing Really Correct Hallucinations?
by: Huang, Baixiang, et al.
Published: (2024)