Med42-v2: A Suite of Clinical LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Christophe, Clément, Kanithi, Praveen K, Raha, Tathagata, Khan, Shadab, Pimentel, Marco AF |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks
by: Pimentel, Marco AF, et al.
Published: (2024)
by: Pimentel, Marco AF, et al.
Published: (2024)
Bridging Language Barriers in Healthcare: A Study on Arabic LLMs
by: Saadi, Nada, et al.
Published: (2025)
by: Saadi, Nada, et al.
Published: (2025)
Named Clinical Entity Recognition Benchmark
by: Abdul, Wadood M, et al.
Published: (2024)
by: Abdul, Wadood M, et al.
Published: (2024)
Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs
by: Christophe, Clément, et al.
Published: (2024)
by: Christophe, Clément, et al.
Published: (2024)
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
by: Kanithi, Praveenkumar, et al.
Published: (2024)
by: Kanithi, Praveenkumar, et al.
Published: (2024)
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
by: Maslenkova, Svetlana, et al.
Published: (2025)
by: Maslenkova, Svetlana, et al.
Published: (2025)
Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches
by: Christophe, Clément, et al.
Published: (2024)
by: Christophe, Clément, et al.
Published: (2024)
Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation
by: Raha, Tathagata, et al.
Published: (2026)
by: Raha, Tathagata, et al.
Published: (2026)
Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare
by: Christophe, Clément, et al.
Published: (2026)
by: Christophe, Clément, et al.
Published: (2026)
Gene42: Long-Range Genomic Foundation Model With Dense Attention
by: Vishniakov, Kirill, et al.
Published: (2025)
by: Vishniakov, Kirill, et al.
Published: (2025)
MedAction: Towards Active Multi-turn Clinical Diagnostic LLMs
by: Hsu, Hsin-Ling, et al.
Published: (2026)
by: Hsu, Hsin-Ling, et al.
Published: (2026)
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
by: Zhang, Ming, et al.
Published: (2025)
by: Zhang, Ming, et al.
Published: (2025)
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
by: Xiao, Yunpeng, et al.
Published: (2025)
by: Xiao, Yunpeng, et al.
Published: (2025)
MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage
by: Khan, Ufaq, et al.
Published: (2026)
by: Khan, Ufaq, et al.
Published: (2026)
MedCT: A Clinical Terminology Graph for Generative AI Applications in Healthcare
by: Chen, Ye, et al.
Published: (2025)
by: Chen, Ye, et al.
Published: (2025)
MedKP: Medical Dialogue with Knowledge Enhancement and Clinical Pathway Encoding
by: Wu, Jiageng, et al.
Published: (2024)
by: Wu, Jiageng, et al.
Published: (2024)
MedOrchestra: A Hybrid Cloud-Local LLM Approach for Clinical Data Interpretation
by: Lee, Sihyeon, et al.
Published: (2025)
by: Lee, Sihyeon, et al.
Published: (2025)
MedReflect: Teaching Medical LLMs to Self-Improve via Reflective Correction
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
by: Xia, Shujun, et al.
Published: (2025)
by: Xia, Shujun, et al.
Published: (2025)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
by: Wei, Jianhui, et al.
Published: (2025)
by: Wei, Jianhui, et al.
Published: (2025)
MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies
by: Ha, Huy Hoang, et al.
Published: (2026)
by: Ha, Huy Hoang, et al.
Published: (2026)
MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical Queries
by: Ghosh, Akash, et al.
Published: (2024)
by: Ghosh, Akash, et al.
Published: (2024)
MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills
by: Yao, Zonghai, et al.
Published: (2024)
by: Yao, Zonghai, et al.
Published: (2024)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
by: Kytöniemi, Joona, et al.
Published: (2025)
by: Kytöniemi, Joona, et al.
Published: (2025)
ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs
by: Ding, Hongxin, et al.
Published: (2025)
by: Ding, Hongxin, et al.
Published: (2025)
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
by: Wu, Juncheng, et al.
Published: (2025)
by: Wu, Juncheng, et al.
Published: (2025)
MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis
by: Chen, Wenting, et al.
Published: (2026)
by: Chen, Wenting, et al.
Published: (2026)
MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication
by: Sambara, Sraavya, et al.
Published: (2026)
by: Sambara, Sraavya, et al.
Published: (2026)
LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
by: He, Ziyuan, et al.
Published: (2025)
by: He, Ziyuan, et al.
Published: (2025)
Can "AI" Be a Doctor? A Study of Empathy, Readability, and Alignment in Clinical LLMs
by: Barone, Mariano, et al.
Published: (2026)
by: Barone, Mariano, et al.
Published: (2026)
MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments
by: Gao, Yicheng, et al.
Published: (2026)
by: Gao, Yicheng, et al.
Published: (2026)
Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs
by: Zun, Lee Qi, et al.
Published: (2025)
by: Zun, Lee Qi, et al.
Published: (2025)
The Zamba2 Suite: Technical Report
by: Glorioso, Paolo, et al.
Published: (2024)
by: Glorioso, Paolo, et al.
Published: (2024)
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
by: Warner, Benjamin, et al.
Published: (2026)
by: Warner, Benjamin, et al.
Published: (2026)
TALES: Text Adventure Learning Environment Suite
by: Cui, Christopher Zhang, et al.
Published: (2025)
by: Cui, Christopher Zhang, et al.
Published: (2025)
M-Prometheus: A Suite of Open Multilingual LLM Judges
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology
by: Zhang, Junlei, et al.
Published: (2023)
by: Zhang, Junlei, et al.
Published: (2023)
Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning
by: Zhang, Xiaotian, et al.
Published: (2025)
by: Zhang, Xiaotian, et al.
Published: (2025)
MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
by: Yang, Lin, et al.
Published: (2026)
by: Yang, Lin, et al.
Published: (2026)
MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
Similar Items
-
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks
by: Pimentel, Marco AF, et al.
Published: (2024) -
Bridging Language Barriers in Healthcare: A Study on Arabic LLMs
by: Saadi, Nada, et al.
Published: (2025) -
Named Clinical Entity Recognition Benchmark
by: Abdul, Wadood M, et al.
Published: (2024) -
Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs
by: Christophe, Clément, et al.
Published: (2024) -
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
by: Kanithi, Praveenkumar, et al.
Published: (2024)