COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
Fuente:
arXiv
Saved in:
| Main Authors: | Panagoulias, Dimitrios P., Papatheodosiou, Persephone, Palamidas, Anastasios P., Sanoudos, Mattheos, Tsoureli-Nikita, Evridiki, Virvou, Maria, Tsihrintzis, George A. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dermacen Analytica: A Novel Methodology Integrating Multi-Modal Large Language Models with Machine Learning in tele-dermatology
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
Modeling Expert AI Diagnostic Alignment via Immutable Inference Snapshots
by: Panagoulias, Dimitrios P., et al.
Published: (2026)
by: Panagoulias, Dimitrios P., et al.
Published: (2026)
Integrating Dynamic Correlation Shifts and Weighted Benchmarking in Extreme Value Analysis
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
Does ejection fraction matter in choosing between percutaneous coronary intervention and coronary artery bypass surgery?
by: Arnold H. Seto, et al.
Published: (2024)
by: Arnold H. Seto, et al.
Published: (2024)
A Sentinel-2 multi-year, multi-country benchmark dataset for crop classification and segmentation with deep learning
by: Sykas, Dimitrios, et al.
Published: (2022)
by: Sykas, Dimitrios, et al.
Published: (2022)
Evaluating the Effectiveness of Training in Improving Nurses' Level of Knowledge and Attitudes Towards Dementia Care in Acute Care Settings: A Mixed‐Method Study
by: Melina Evripidou, et al.
Published: (2026)
by: Melina Evripidou, et al.
Published: (2026)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
by: Montalan, Jann Railey, et al.
Published: (2025)
by: Montalan, Jann Railey, et al.
Published: (2025)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
by: Li, Kunning, et al.
Published: (2025)
by: Li, Kunning, et al.
Published: (2025)
BAAF: A benchmark attention adaptive framework for medical ultrasound image segmentation tasks
by: Chen, Gongping, et al.
Published: (2023)
by: Chen, Gongping, et al.
Published: (2023)
Raygun benchmarking + Training dataset
by: Devkota, Kapil
Published: (2026)
by: Devkota, Kapil
Published: (2026)
Toward a benchmark for CTR prediction in online advertising: datasets, evaluation protocols and perspectives
by: Gao, Shan, et al.
Published: (2025)
by: Gao, Shan, et al.
Published: (2025)
Extreme Weather Bench: A framework and benchmark for evaluation of high-impact weather
by: McGovern, Amy, et al.
Published: (2026)
by: McGovern, Amy, et al.
Published: (2026)
Qwen-BIM: developing large language model for BIM-based design with domain-specific benchmark and dataset
by: Lin, Jia-Rui, et al.
Published: (2026)
by: Lin, Jia-Rui, et al.
Published: (2026)
A Self supervised learning framework for imbalanced medical imaging datasets
by: Sharma, Yash Kumar, et al.
Published: (2026)
by: Sharma, Yash Kumar, et al.
Published: (2026)
Changes in soil test phosphorus and soil cations following application of sewage sludge ash and other recycled phosphorus fertilizers
by: Persephone Ma, et al.
Published: (2025)
by: Persephone Ma, et al.
Published: (2025)
Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Language corpora for the Dutch medical domain
by: van Es, B.
Published: (2026)
by: van Es, B.
Published: (2026)
Fine-Grained Customized Fashion Design with Image-into-Prompt benchmark and dataset from LMM
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Motor and Non‐motor Outcomes of Deep Brain Stimulation across the Genetic Panorama of Parkinson's Disease: A Multi‐Scale Meta‐Analysis
by: Evridiki Asimakidou, et al.
Published: (2024)
by: Evridiki Asimakidou, et al.
Published: (2024)
A comprehensive and easy-to-use multi-domain multi-task medical imaging meta-dataset
by: Woerner, Stefano, et al.
Published: (2024)
by: Woerner, Stefano, et al.
Published: (2024)
Datasheets for AI and medical datasets (DAIMS): a data validation and documentation framework before machine learning analysis in medical research
by: Marandi, Ramtin Zargari, et al.
Published: (2025)
by: Marandi, Ramtin Zargari, et al.
Published: (2025)
3DBubbles : An experimental dataset for model training and benchmarking
by: Baodi Yu, et al.
Published: (2025)
by: Baodi Yu, et al.
Published: (2025)
Harnessing Large Language Model to collect and analyze Metal-organic framework property dataset
by: Lee, Wonseok, et al.
Published: (2024)
by: Lee, Wonseok, et al.
Published: (2024)
Hierarchical Pattern Decryption Methodology for Ransomware Detection Using Probabilistic Cryptographic Footprints
by: Pekepok, Kevin, et al.
Published: (2025)
by: Pekepok, Kevin, et al.
Published: (2025)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
by: Rädsch, Tim, et al.
Published: (2025)
by: Rädsch, Tim, et al.
Published: (2025)
Uniformization of Gromov hyperbolic domains by circle domains
by: Karafyllia, Christina, et al.
Published: (2024)
by: Karafyllia, Christina, et al.
Published: (2024)
SAM2CLIP2SAM: Vision Language Model for Segmentation of 3D CT Scans for Covid-19 Detection
by: Kollias, Dimitrios, et al.
Published: (2024)
by: Kollias, Dimitrios, et al.
Published: (2024)
A framework for medical physics compensation in an academic department
by: David P. Gierga, et al.
Published: (2024)
by: David P. Gierga, et al.
Published: (2024)
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
by: Ging, Simon, et al.
Published: (2024)
by: Ging, Simon, et al.
Published: (2024)
Large language models in healthcare and medical domain: A review
by: Nazi, Zabir Al, et al.
Published: (2023)
by: Nazi, Zabir Al, et al.
Published: (2023)
Authors reply: IL‐1β/DNA complex elevation distinguishes autoinflammatory disorders from autoimmune and infectious diseases
by: Anastasia‐Maria Natsi, et al.
Published: (2024)
by: Anastasia‐Maria Natsi, et al.
Published: (2024)
DistillER: Knowledge Distillation in Entity Resolution with Large Language Models
by: Zeakis, Alexandros, et al.
Published: (2026)
by: Zeakis, Alexandros, et al.
Published: (2026)
FLOGA: A machine learning ready dataset, a benchmark and a novel deep learning model for burnt area mapping with Sentinel-2
by: Sdraka, Maria, et al.
Published: (2023)
by: Sdraka, Maria, et al.
Published: (2023)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Staining normalization in histopathology: Method benchmarking using multicenter dataset
by: Khan, Umair, et al.
Published: (2025)
by: Khan, Umair, et al.
Published: (2025)
Cueless EEG imagined speech for subject identification: dataset and benchmarks
by: Derakhshesh, Ali, et al.
Published: (2025)
by: Derakhshesh, Ali, et al.
Published: (2025)
A large dataset curation and benchmark for drug target interaction
by: Golts, Alex, et al.
Published: (2024)
by: Golts, Alex, et al.
Published: (2024)
How well do LLMs cite relevant medical references? An evaluation framework and analyses
by: Wu, Kevin, et al.
Published: (2024)
by: Wu, Kevin, et al.
Published: (2024)
What do model reports say about their ChemBio benchmark evaluations? Comparing recent releases to the STREAM framework
by: Reed, Tom, et al.
Published: (2025)
by: Reed, Tom, et al.
Published: (2025)
Similar Items
-
Dermacen Analytica: A Novel Methodology Integrating Multi-Modal Large Language Models with Machine Learning in tele-dermatology
by: Panagoulias, Dimitrios P., et al.
Published: (2024) -
Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis
by: Panagoulias, Dimitrios P., et al.
Published: (2024) -
Modeling Expert AI Diagnostic Alignment via Immutable Inference Snapshots
by: Panagoulias, Dimitrios P., et al.
Published: (2026) -
Integrating Dynamic Correlation Shifts and Weighted Benchmarking in Extreme Value Analysis
by: Panagoulias, Dimitrios P., et al.
Published: (2024) -
Does ejection fraction matter in choosing between percutaneous coronary intervention and coronary artery bypass surgery?
by: Arnold H. Seto, et al.
Published: (2024)