Comparative Experimentation of Accuracy Metrics in Automated Medical Reporting: The Case of Otitis Consultations
Fuente:
arXiv
Guardado en:
| Autores principales: | Faber, Wouter, Bootsma, Renske Eline, Huibers, Tom, van Dulmen, Sandra, Brinkkemper, Sjaak |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Enhancing Summarization Performance through Transformer-Based Prompt Engineering in Automated Medical Reporting
por: van Zandvoort, Daphne, et al.
Publicado: (2023)
por: van Zandvoort, Daphne, et al.
Publicado: (2023)
Summarizing long regulatory documents with a multi-step pipeline
por: Sie, Mika, et al.
Publicado: (2024)
por: Sie, Mika, et al.
Publicado: (2024)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
por: Kocmi, Tom, et al.
Publicado: (2024)
por: Kocmi, Tom, et al.
Publicado: (2024)
GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training
por: van Oort, Jesse, et al.
Publicado: (2026)
por: van Oort, Jesse, et al.
Publicado: (2026)
Healthcare Copilot: Eliciting the Power of General LLMs for Medical Consultation
por: Ren, Zhiyao, et al.
Publicado: (2024)
por: Ren, Zhiyao, et al.
Publicado: (2024)
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
por: Deviyani, Athiya, et al.
Publicado: (2025)
por: Deviyani, Athiya, et al.
Publicado: (2025)
Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation
por: Ren, Zhiyao, et al.
Publicado: (2026)
por: Ren, Zhiyao, et al.
Publicado: (2026)
Comparing Large Language Models and Traditional Machine Translation Tools for Translating Medical Consultation Summaries: A Pilot Study
por: Li, Andy, et al.
Publicado: (2025)
por: Li, Andy, et al.
Publicado: (2025)
Beyond Idealized Patients: Evaluating LLMs under Challenging Patient Behaviors in Medical Consultations
por: Li, Yahan, et al.
Publicado: (2026)
por: Li, Yahan, et al.
Publicado: (2026)
Tool Calling: Enhancing Medication Consultation via Retrieval-Augmented Large Language Models
por: Huang, Zhongzhen, et al.
Publicado: (2024)
por: Huang, Zhongzhen, et al.
Publicado: (2024)
EMRModel: A Large Language Model for Extracting Medical Consultation Dialogues into Structured Medical Records
por: Zhao, Shuguang, et al.
Publicado: (2025)
por: Zhao, Shuguang, et al.
Publicado: (2025)
Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation
por: Lim, Seungseop, et al.
Publicado: (2025)
por: Lim, Seungseop, et al.
Publicado: (2025)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
por: Cho, Yousang, et al.
Publicado: (2025)
por: Cho, Yousang, et al.
Publicado: (2025)
ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical Judgment
por: Li, Ruochen, et al.
Publicado: (2025)
por: Li, Ruochen, et al.
Publicado: (2025)
CASPR: Automated Evaluation Metric for Contrastive Summarization
por: Ananthamurugan, Nirupan, et al.
Publicado: (2024)
por: Ananthamurugan, Nirupan, et al.
Publicado: (2024)
An Automated Length-Aware Quality Metric for Summarization
por: Foland, Andrew D.
Publicado: (2025)
por: Foland, Andrew D.
Publicado: (2025)
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
por: Chen, Sirry, et al.
Publicado: (2026)
por: Chen, Sirry, et al.
Publicado: (2026)
Satisfactory Medical Consultation based on Terminology-Enhanced Information Retrieval and Emotional In-Context Learning
por: Zuo, Kaiwen, et al.
Publicado: (2025)
por: Zuo, Kaiwen, et al.
Publicado: (2025)
Research on Medical Named Entity Identification Based On Prompt-Biomrc Model and Its Application in Intelligent Consultation System
por: Yang, Jinzhu
Publicado: (2025)
por: Yang, Jinzhu
Publicado: (2025)
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
por: Liu, Hongcheng, et al.
Publicado: (2024)
por: Liu, Hongcheng, et al.
Publicado: (2024)
A Debate-Driven Experiment on LLM Hallucinations and Accuracy
por: Li, Ray, et al.
Publicado: (2024)
por: Li, Ray, et al.
Publicado: (2024)
Automated Medical Report Generation for ECG Data: Bridging Medical Text and Signal Processing with Deep Learning
por: Bleich, Amnon, et al.
Publicado: (2024)
por: Bleich, Amnon, et al.
Publicado: (2024)
Yes, this is what I was looking for! Towards Multi-modal Medical Consultation Concern Summary Generation
por: Tiwari, Abhisek, et al.
Publicado: (2024)
por: Tiwari, Abhisek, et al.
Publicado: (2024)
Granular Change Accuracy: A More Accurate Performance Metric for Dialogue State Tracking
por: Aksu, Taha, et al.
Publicado: (2024)
por: Aksu, Taha, et al.
Publicado: (2024)
Claim Extraction for Fact-Checking: Data, Models, and Automated Metrics
por: Ullrich, Herbert, et al.
Publicado: (2025)
por: Ullrich, Herbert, et al.
Publicado: (2025)
Overview of the MEDIQA-OE 2025 Shared Task on Medical Order Extraction from Doctor-Patient Consultations
por: Corbeil, Jean-Philippe, et al.
Publicado: (2025)
por: Corbeil, Jean-Philippe, et al.
Publicado: (2025)
A Measure of the System Dependence of Automated Metrics
por: von Däniken, Pius, et al.
Publicado: (2024)
por: von Däniken, Pius, et al.
Publicado: (2024)
The Curious Case of Representational Alignment: Unravelling Visio-Linguistic Tasks in Emergent Communication
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
por: Kouwenhoven, Tom, et al.
Publicado: (2024)
LexGenie: Automated Generation of Structured Reports for European Court of Human Rights Case Law
por: Santosh, T. Y. S. S, et al.
Publicado: (2025)
por: Santosh, T. Y. S. S, et al.
Publicado: (2025)
Voice Under Revision: Large Language Models and the Normalization of Personal Narrative
por: van Nuenen, Tom
Publicado: (2026)
por: van Nuenen, Tom
Publicado: (2026)
Recognition Without Authorization: LLMs and the Moral Order of Online Advice
por: van Nuenen, Tom
Publicado: (2026)
por: van Nuenen, Tom
Publicado: (2026)
Unveiling the Tapestry of Automated Essay Scoring: A Comprehensive Investigation of Accuracy, Fairness, and Generalizability
por: Yang, Kaixun, et al.
Publicado: (2024)
por: Yang, Kaixun, et al.
Publicado: (2024)
DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent Collaboration
por: Jia, Zhihao, et al.
Publicado: (2025)
por: Jia, Zhihao, et al.
Publicado: (2025)
Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
por: Doshi, Savan
Publicado: (2026)
por: Doshi, Savan
Publicado: (2026)
Consultation on Industrial Machine Faults with Large language Models
por: Boonmee, Apiradee, et al.
Publicado: (2024)
por: Boonmee, Apiradee, et al.
Publicado: (2024)
Comparing Hallucination Detection Metrics for Multilingual Generation
por: Kang, Haoqiang, et al.
Publicado: (2024)
por: Kang, Haoqiang, et al.
Publicado: (2024)
Compared to What? Baselines and Metrics for Counterfactual Prompting
por: Yang, Zihao, et al.
Publicado: (2026)
por: Yang, Zihao, et al.
Publicado: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
por: Cavalin, Paulo, et al.
Publicado: (2025)
por: Cavalin, Paulo, et al.
Publicado: (2025)
RaTEScore: A Metric for Radiology Report Generation
por: Zhao, Weike, et al.
Publicado: (2024)
por: Zhao, Weike, et al.
Publicado: (2024)
Ejemplares similares
-
Enhancing Summarization Performance through Transformer-Based Prompt Engineering in Automated Medical Reporting
por: van Zandvoort, Daphne, et al.
Publicado: (2023) -
Summarizing long regulatory documents with a multi-step pipeline
por: Sie, Mika, et al.
Publicado: (2024) -
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
por: Kocmi, Tom, et al.
Publicado: (2024) -
GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training
por: van Oort, Jesse, et al.
Publicado: (2026) -
Healthcare Copilot: Eliciting the Power of General LLMs for Medical Consultation
por: Ren, Zhiyao, et al.
Publicado: (2024)