A comprehensive multimodal dataset and benchmark for ulcerative colitis scoring in endoscopy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghatwary, Noha, Yue, Jiangbei, Elgendy, Ahmed, Nagdy, Hanna, Galal, Ahmed, Fathy, Hayam, El-Amin, Hussein, Subramanian, Venkataraman, Mohammed, Noor, Ochoa-Ruiz, Gilberto, Ali, Sharib
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912967362084864
author Ghatwary, Noha
Yue, Jiangbei
Elgendy, Ahmed
Nagdy, Hanna
Galal, Ahmed
Fathy, Hayam
El-Amin, Hussein
Subramanian, Venkataraman
Mohammed, Noor
Ochoa-Ruiz, Gilberto
Ali, Sharib
author_facet Ghatwary, Noha
Yue, Jiangbei
Elgendy, Ahmed
Nagdy, Hanna
Galal, Ahmed
Fathy, Hayam
El-Amin, Hussein
Subramanian, Venkataraman
Mohammed, Noor
Ochoa-Ruiz, Gilberto
Ali, Sharib
contents Ulcerative colitis (UC) is a chronic mucosal inflammatory condition that places patients at increased risk of colorectal cancer. Colonoscopic surveillance remains the gold standard for assessing disease activity, and reporting typically relies on standardised endoscopic scoring metrics. The most widely used is the Mayo Endoscopic Score (MES), with some centres also adopting the Ulcerative Colitis Endoscopic Index of Severity (UCEIS). Both are descriptive assessments of mucosal inflammation (MES: 0 to 3; UCEIS: 0 to 8), where higher values indicate more severe disease. However, computational methods for automatically predicting these scores remain limited, largely due to the lack of publicly available expert-annotated datasets and the absence of robust benchmarking. There is also a significant research gap in generating clinically meaningful descriptions of UC images, despite image captioning being a well-established computer vision task. Variability in endoscopic systems and procedural workflows across centres further highlights the need for multi-centre datasets to ensure algorithmic robustness and generalisability. In this work, we introduce a curated multi-centre, multi-resolution dataset that includes expert-validated MES and UCEIS labels, alongside detailed clinical descriptions. To our knowledge, this is the first comprehensive dataset that combines dual scoring metrics for classification tasks with expert-generated captions describing mucosal appearance and clinically accepted reasoning for image captioning. This resource opens new opportunities for developing clinically meaningful multimodal algorithms. In addition to the dataset, we also provide benchmarking using convolutional neural networks, vision transformers, hybrid models, and widely used multimodal vision-language captioning algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14559
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A comprehensive multimodal dataset and benchmark for ulcerative colitis scoring in endoscopy
Ghatwary, Noha
Yue, Jiangbei
Elgendy, Ahmed
Nagdy, Hanna
Galal, Ahmed
Fathy, Hayam
El-Amin, Hussein
Subramanian, Venkataraman
Mohammed, Noor
Ochoa-Ruiz, Gilberto
Ali, Sharib
Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
Ulcerative colitis (UC) is a chronic mucosal inflammatory condition that places patients at increased risk of colorectal cancer. Colonoscopic surveillance remains the gold standard for assessing disease activity, and reporting typically relies on standardised endoscopic scoring metrics. The most widely used is the Mayo Endoscopic Score (MES), with some centres also adopting the Ulcerative Colitis Endoscopic Index of Severity (UCEIS). Both are descriptive assessments of mucosal inflammation (MES: 0 to 3; UCEIS: 0 to 8), where higher values indicate more severe disease. However, computational methods for automatically predicting these scores remain limited, largely due to the lack of publicly available expert-annotated datasets and the absence of robust benchmarking. There is also a significant research gap in generating clinically meaningful descriptions of UC images, despite image captioning being a well-established computer vision task. Variability in endoscopic systems and procedural workflows across centres further highlights the need for multi-centre datasets to ensure algorithmic robustness and generalisability. In this work, we introduce a curated multi-centre, multi-resolution dataset that includes expert-validated MES and UCEIS labels, alongside detailed clinical descriptions. To our knowledge, this is the first comprehensive dataset that combines dual scoring metrics for classification tasks with expert-generated captions describing mucosal appearance and clinically accepted reasoning for image captioning. This resource opens new opportunities for developing clinically meaningful multimodal algorithms. In addition to the dataset, we also provide benchmarking using convolutional neural networks, vision transformers, hybrid models, and widely used multimodal vision-language captioning algorithms.
title A comprehensive multimodal dataset and benchmark for ulcerative colitis scoring in endoscopy
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2603.14559