Medico 2025: Visual Question Answering for Gastrointestinal Imaging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gautam, Sushant, Thambawita, Vajira, Riegler, Michael, Halvorsen, Pål, Hicks, Steven
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915446832234496
author Gautam, Sushant
Thambawita, Vajira
Riegler, Michael
Halvorsen, Pål
Hicks, Steven
author_facet Gautam, Sushant
Thambawita, Vajira
Riegler, Michael
Halvorsen, Pål
Hicks, Steven
contents The Medico 2025 challenge addresses Visual Question Answering (VQA) for Gastrointestinal (GI) imaging, organized as part of the MediaEval task series. The challenge focuses on developing Explainable Artificial Intelligence (XAI) models that answer clinically relevant questions based on GI endoscopy images while providing interpretable justifications aligned with medical reasoning. It introduces two subtasks: (1) answering diverse types of visual questions using the Kvasir-VQA-x1 dataset, and (2) generating multimodal explanations to support clinical decision-making. The Kvasir-VQA-x1 dataset, created from 6,500 images and 159,549 complex question-answer (QA) pairs, serves as the benchmark for the challenge. By combining quantitative performance metrics and expert-reviewed explainability assessments, this task aims to advance trustworthy Artificial Intelligence (AI) in medical image analysis. Instructions, data access, and an updated guide for participation are available in the official competition repository: https://github.com/simula/MediaEval-Medico-2025
format Preprint
id arxiv_https___arxiv_org_abs_2508_10869
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Medico 2025: Visual Question Answering for Gastrointestinal Imaging
Gautam, Sushant
Thambawita, Vajira
Riegler, Michael
Halvorsen, Pål
Hicks, Steven
Computer Vision and Pattern Recognition
Artificial Intelligence
68T45, 92C55
I.2.10; I.4.9
The Medico 2025 challenge addresses Visual Question Answering (VQA) for Gastrointestinal (GI) imaging, organized as part of the MediaEval task series. The challenge focuses on developing Explainable Artificial Intelligence (XAI) models that answer clinically relevant questions based on GI endoscopy images while providing interpretable justifications aligned with medical reasoning. It introduces two subtasks: (1) answering diverse types of visual questions using the Kvasir-VQA-x1 dataset, and (2) generating multimodal explanations to support clinical decision-making. The Kvasir-VQA-x1 dataset, created from 6,500 images and 159,549 complex question-answer (QA) pairs, serves as the benchmark for the challenge. By combining quantitative performance metrics and expert-reviewed explainability assessments, this task aims to advance trustworthy Artificial Intelligence (AI) in medical image analysis. Instructions, data access, and an updated guide for participation are available in the official competition repository: https://github.com/simula/MediaEval-Medico-2025
title Medico 2025: Visual Question Answering for Gastrointestinal Imaging
topic Computer Vision and Pattern Recognition
Artificial Intelligence
68T45, 92C55
I.2.10; I.4.9
url https://arxiv.org/abs/2508.10869