MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
Fuente:
arXiv
Guardado en:
| Autores principales: | Abacha, Asma Ben, Yim, Wen-wai, Fu, Yujuan, Sun, Zhaoyi, Yetisgen, Meliha, Xia, Fei, Lin, Thomas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
RADAR: A Multimodal Benchmark for 3D Image-Based Radiology Report Review
por: Sun, Zhaoyi, et al.
Publicado: (2026)
por: Sun, Zhaoyi, et al.
Publicado: (2026)
DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
por: Sun, Zhaoyi, et al.
Publicado: (2025)
por: Sun, Zhaoyi, et al.
Publicado: (2025)
Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions
por: Fu, Yujuan, et al.
Publicado: (2024)
por: Fu, Yujuan, et al.
Publicado: (2024)
A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction
por: Ahmed, Farzad, et al.
Publicado: (2025)
por: Ahmed, Farzad, et al.
Publicado: (2025)
BioMistral-NLU: Towards More Generalizable Medical Language Understanding through Instruction Tuning
por: Fu, Yujuan Velvin, et al.
Publicado: (2024)
por: Fu, Yujuan Velvin, et al.
Publicado: (2024)
CACER: Clinical Concept Annotations for Cancer Events and Relations
por: Fu, Yujuan, et al.
Publicado: (2024)
por: Fu, Yujuan, et al.
Publicado: (2024)
Extracting Social Determinants of Health from Pediatric Patient Notes Using Large Language Models: Novel Corpus and Methods
por: Fu, Yujuan, et al.
Publicado: (2024)
por: Fu, Yujuan, et al.
Publicado: (2024)
RadTimeline: Timeline Summarization for Longitudinal Radiological Lung Findings
por: Zhou, Sitong, et al.
Publicado: (2026)
por: Zhou, Sitong, et al.
Publicado: (2026)
UW-BioNLP at ChemoTimelines 2025: Thinking, Fine-Tuning, and Dictionary-Enhanced LLM Systems for Chemotherapy Timeline Extraction
por: Zhang, Tianmai M., et al.
Publicado: (2025)
por: Zhang, Tianmai M., et al.
Publicado: (2025)
Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches
por: Park, Namu, et al.
Publicado: (2025)
por: Park, Namu, et al.
Publicado: (2025)
Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
por: Dobbins, Nic, et al.
Publicado: (2025)
por: Dobbins, Nic, et al.
Publicado: (2025)
Identifying Imaging Follow-Up in Radiology Reports: A Comparative Analysis of Traditional ML and LLM Approaches
por: Park, Namu, et al.
Publicado: (2025)
por: Park, Namu, et al.
Publicado: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
por: Bologna, Federica, et al.
Publicado: (2026)
por: Bologna, Federica, et al.
Publicado: (2026)
Adapting Biomedical Abstracts into Plain language using Large Language Models
por: Gangavarapu, Haritha, et al.
Publicado: (2025)
por: Gangavarapu, Haritha, et al.
Publicado: (2025)
A Novel Corpus of Annotated Medical Imaging Reports and Information Extraction Results Using BERT-based Language Models
por: Park, Namu, et al.
Publicado: (2024)
por: Park, Namu, et al.
Publicado: (2024)
CoRe-BT: A Multimodal Radiology-Pathology-Text Benchmark for Robust Brain Tumor Typing
por: Rivera, Juampablo E. Heras, et al.
Publicado: (2026)
por: Rivera, Juampablo E. Heras, et al.
Publicado: (2026)
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
por: Sakib, Fardin Ahsan, et al.
Publicado: (2025)
por: Sakib, Fardin Ahsan, et al.
Publicado: (2025)
MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations
por: Ma, Congbo, et al.
Publicado: (2026)
por: Ma, Congbo, et al.
Publicado: (2026)
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
por: Corbeil, Jean-Philippe, et al.
Publicado: (2025)
por: Corbeil, Jean-Philippe, et al.
Publicado: (2025)
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
por: Iwase, Naoto, et al.
Publicado: (2025)
por: Iwase, Naoto, et al.
Publicado: (2025)
Overview of the MEDIQA-OE 2025 Shared Task on Medical Order Extraction from Doctor-Patient Consultations
por: Corbeil, Jean-Philippe, et al.
Publicado: (2025)
por: Corbeil, Jean-Philippe, et al.
Publicado: (2025)
Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication
por: Chung, Philip, et al.
Publicado: (2024)
por: Chung, Philip, et al.
Publicado: (2024)
Importance of Prompt Optimisation for Error Detection in Medical Notes Using Language Models
por: Myles, Craig, et al.
Publicado: (2026)
por: Myles, Craig, et al.
Publicado: (2026)
Term2Note: Synthesising Differentially Private Clinical Notes from Medical Terms
por: Wu, Yuping, et al.
Publicado: (2025)
por: Wu, Yuping, et al.
Publicado: (2025)
AR-BENCH: Benchmarking Legal Reasoning with Judgment Error Detection, Classification and Correction
por: Li, Yifei, et al.
Publicado: (2026)
por: Li, Yifei, et al.
Publicado: (2026)
Traj-CoA: Patient Trajectory Modeling via Chain-of-Agents for Lung Cancer Risk Prediction
por: Zeng, Sihang, et al.
Publicado: (2025)
por: Zeng, Sihang, et al.
Publicado: (2025)
Empowering Healthcare Practitioners with Language Models: Structuring Speech Transcripts in Two Real-World Clinical Applications
por: Corbeil, Jean-Philippe, et al.
Publicado: (2025)
por: Corbeil, Jean-Philippe, et al.
Publicado: (2025)
WangLab at MEDIQA-CORR 2024: Optimized LLM-based Programs for Medical Error Detection and Correction
por: Toma, Augustin, et al.
Publicado: (2024)
por: Toma, Augustin, et al.
Publicado: (2024)
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
por: Byun, Ji Young, et al.
Publicado: (2026)
por: Byun, Ji Young, et al.
Publicado: (2026)
IryoNLP at MEDIQA-CORR 2024: Tackling the Medical Error Detection & Correction Task On the Shoulders of Medical Agents
por: Corbeil, Jean-Philippe
Publicado: (2024)
por: Corbeil, Jean-Philippe
Publicado: (2024)
Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical Notes
por: Zhou, Yang, et al.
Publicado: (2026)
por: Zhou, Yang, et al.
Publicado: (2026)
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
por: Yan, Yibo, et al.
Publicado: (2024)
por: Yan, Yibo, et al.
Publicado: (2024)
Detection-Correction Structure via General Language Model for Grammatical Error Correction
por: Li, Wei, et al.
Publicado: (2024)
por: Li, Wei, et al.
Publicado: (2024)
The Task-oriented Queries Benchmark (ToQB)
por: Yim, Keun Soo
Publicado: (2024)
por: Yim, Keun Soo
Publicado: (2024)
EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error Correction
por: Ye, Jingheng, et al.
Publicado: (2024)
por: Ye, Jingheng, et al.
Publicado: (2024)
Racing Thoughts: Explaining Contextualization Errors in Large Language Models
por: Lepori, Michael A., et al.
Publicado: (2024)
por: Lepori, Michael A., et al.
Publicado: (2024)
Benchmarking Multi-turn Medical Diagnosis: Hold, Lure, and Self-Correction
por: Fang, Jinrui, et al.
Publicado: (2026)
por: Fang, Jinrui, et al.
Publicado: (2026)
Personalized Clinical Note Generation from Doctor-Patient Conversations
por: Brake, Nathan, et al.
Publicado: (2024)
por: Brake, Nathan, et al.
Publicado: (2024)
Ejemplares similares
-
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025) -
RADAR: A Multimodal Benchmark for 3D Image-Based Radiology Report Review
por: Sun, Zhaoyi, et al.
Publicado: (2026) -
DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images
por: Yim, Wen-wai, et al.
Publicado: (2025) -
A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
por: Sun, Zhaoyi, et al.
Publicado: (2025) -
Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions
por: Fu, Yujuan, et al.
Publicado: (2024)