Frontiers in Intelligent Colonoscopy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Ge-Peng, Liu, Jingyi, Xu, Peng, Barnes, Nick, Khan, Fahad Shahbaz, Khan, Salman, Fan, Deng-Ping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912601937543168
author Ji, Ge-Peng
Liu, Jingyi
Xu, Peng
Barnes, Nick
Khan, Fahad Shahbaz
Khan, Salman
Fan, Deng-Ping
author_facet Ji, Ge-Peng
Liu, Jingyi
Xu, Peng
Barnes, Nick
Khan, Fahad Shahbaz
Khan, Salman
Fan, Deng-Ping
contents Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal medical applications. With this goal, we begin by assessing the current data-centric and model-centric landscapes through four tasks for colonoscopic scene perception, including classification, detection, segmentation, and vision-language understanding. This assessment enables us to identify domain-specific challenges and reveals that multimodal research in colonoscopy remains open for further exploration. To embrace the coming multimodal era, we establish three foundational initiatives: a large-scale multimodal instruction tuning dataset ColonINST, a colonoscopy-designed multimodal language model ColonGPT, and a multimodal benchmark. To facilitate ongoing monitoring of this rapidly evolving field, we provide a public website for the latest updates: https://github.com/ai4colonoscopy/IntelliScope.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17241
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Frontiers in Intelligent Colonoscopy
Ji, Ge-Peng
Liu, Jingyi
Xu, Peng
Barnes, Nick
Khan, Fahad Shahbaz
Khan, Salman
Fan, Deng-Ping
Image and Video Processing
Computer Vision and Pattern Recognition
Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal medical applications. With this goal, we begin by assessing the current data-centric and model-centric landscapes through four tasks for colonoscopic scene perception, including classification, detection, segmentation, and vision-language understanding. This assessment enables us to identify domain-specific challenges and reveals that multimodal research in colonoscopy remains open for further exploration. To embrace the coming multimodal era, we establish three foundational initiatives: a large-scale multimodal instruction tuning dataset ColonINST, a colonoscopy-designed multimodal language model ColonGPT, and a multimodal benchmark. To facilitate ongoing monitoring of this rapidly evolving field, we provide a public website for the latest updates: https://github.com/ai4colonoscopy/IntelliScope.
title Frontiers in Intelligent Colonoscopy
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.17241