Medical AI Consensus: A Multi-Agent Framework for Radiology Report Generation and Evaluation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Elboardy, Ahmed T., Khoriba, Ghada, Rashed, Essam A.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912598871506944
author Elboardy, Ahmed T.
Khoriba, Ghada
Rashed, Essam A.
author_facet Elboardy, Ahmed T.
Khoriba, Ghada
Rashed, Essam A.
contents Automating radiology report generation poses a dual challenge: building clinically reliable systems and designing rigorous evaluation protocols. We introduce a multi-agent reinforcement learning framework that serves as both a benchmark and evaluation environment for multimodal clinical reasoning in the radiology ecosystem. The proposed framework integrates large language models (LLMs) and large vision models (LVMs) within a modular architecture composed of ten specialized agents responsible for image analysis, feature extraction, report generation, review, and evaluation. This design enables fine-grained assessment at both the agent level (e.g., detection and segmentation accuracy) and the consensus level (e.g., report quality and clinical relevance). We demonstrate an implementation using chatGPT-4o on public radiology datasets, where LLMs act as evaluators alongside medical radiologist feedback. By aligning evaluation protocols with the LLM development lifecycle, including pretraining, finetuning, alignment, and deployment, the proposed benchmark establishes a path toward trustworthy deviance-based radiology report generation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17353
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Medical AI Consensus: A Multi-Agent Framework for Radiology Report Generation and Evaluation
Elboardy, Ahmed T.
Khoriba, Ghada
Rashed, Essam A.
Artificial Intelligence
Image and Video Processing
Medical Physics
Automating radiology report generation poses a dual challenge: building clinically reliable systems and designing rigorous evaluation protocols. We introduce a multi-agent reinforcement learning framework that serves as both a benchmark and evaluation environment for multimodal clinical reasoning in the radiology ecosystem. The proposed framework integrates large language models (LLMs) and large vision models (LVMs) within a modular architecture composed of ten specialized agents responsible for image analysis, feature extraction, report generation, review, and evaluation. This design enables fine-grained assessment at both the agent level (e.g., detection and segmentation accuracy) and the consensus level (e.g., report quality and clinical relevance). We demonstrate an implementation using chatGPT-4o on public radiology datasets, where LLMs act as evaluators alongside medical radiologist feedback. By aligning evaluation protocols with the LLM development lifecycle, including pretraining, finetuning, alignment, and deployment, the proposed benchmark establishes a path toward trustworthy deviance-based radiology report generation.
title Medical AI Consensus: A Multi-Agent Framework for Radiology Report Generation and Evaluation
topic Artificial Intelligence
Image and Video Processing
Medical Physics
url https://arxiv.org/abs/2509.17353