3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918195001032704 |
|---|---|
| author | Sviridov, Ivan Miftakhova, Amina Tereshchenko, Artemiy Zubkova, Galina Blinov, Pavel Savchenko, Andrey |
| author_facet | Sviridov, Ivan Miftakhova, Amina Tereshchenko, Artemiy Zubkova, Galina Blinov, Pavel Savchenko, Andrey |
| contents | Though Large Vision-Language Models (LVLMs) are being actively explored in medicine, their ability to conduct complex real-world telemedicine consultations combining accurate diagnosis with professional dialogue remains underexplored. This paper presents 3MDBench (Medical Multimodal Multi-agent Dialogue Benchmark), an open-source framework for simulating and evaluating LVLM-driven telemedical consultations. 3MDBench simulates patient variability through temperament-based Patient Agent and evaluates diagnostic accuracy and dialogue quality via Assessor Agent. It includes 2996 cases across 34 diagnoses from real-world telemedicine interactions, combining textual and image-based data. The experimental study compares diagnostic strategies for widely used open and closed-source LVLMs. We demonstrate that multimodal dialogue with internal reasoning improves F1 score by 6.5% over non-dialogue settings, highlighting the importance of context-aware, information-seeking questioning. Moreover, injecting predictions from a diagnostic convolutional neural network into the LVLM's context boosts F1 by up to 20%. Source code is available at https://github.com/univanxx/3mdbench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_13861 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | 3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark Sviridov, Ivan Miftakhova, Amina Tereshchenko, Artemiy Zubkova, Galina Blinov, Pavel Savchenko, Andrey Human-Computer Interaction Computation and Language Multiagent Systems 68T42 I.2.1 Though Large Vision-Language Models (LVLMs) are being actively explored in medicine, their ability to conduct complex real-world telemedicine consultations combining accurate diagnosis with professional dialogue remains underexplored. This paper presents 3MDBench (Medical Multimodal Multi-agent Dialogue Benchmark), an open-source framework for simulating and evaluating LVLM-driven telemedical consultations. 3MDBench simulates patient variability through temperament-based Patient Agent and evaluates diagnostic accuracy and dialogue quality via Assessor Agent. It includes 2996 cases across 34 diagnoses from real-world telemedicine interactions, combining textual and image-based data. The experimental study compares diagnostic strategies for widely used open and closed-source LVLMs. We demonstrate that multimodal dialogue with internal reasoning improves F1 score by 6.5% over non-dialogue settings, highlighting the importance of context-aware, information-seeking questioning. Moreover, injecting predictions from a diagnostic convolutional neural network into the LVLM's context boosts F1 by up to 20%. Source code is available at https://github.com/univanxx/3mdbench. |
| title | 3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark |
| topic | Human-Computer Interaction Computation and Language Multiagent Systems 68T42 I.2.1 |
| url | https://arxiv.org/abs/2504.13861 |