AI-Assisted Colonoscopy: Polyp Detection and Segmentation using Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Delaquintana-Aramendi, Uxue, Benito-del-Valle, Leire, Alvarez-Gila, Aitor, Pascau, Javier, Sánchez-Peralta, Luisa F, Picón, Artzai, Pagador, J Blas, Saratxaga, Cristina L
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915219322699776
author Delaquintana-Aramendi, Uxue
Benito-del-Valle, Leire
Alvarez-Gila, Aitor
Pascau, Javier
Sánchez-Peralta, Luisa F
Picón, Artzai
Pagador, J Blas
Saratxaga, Cristina L
author_facet Delaquintana-Aramendi, Uxue
Benito-del-Valle, Leire
Alvarez-Gila, Aitor
Pascau, Javier
Sánchez-Peralta, Luisa F
Picón, Artzai
Pagador, J Blas
Saratxaga, Cristina L
contents In colonoscopy, 80% of the missed polyps could be detected with the help of Deep Learning models. In the search for algorithms capable of addressing this challenge, foundation models emerge as promising candidates. Their zero-shot or few-shot learning capabilities, facilitate generalization to new data or tasks without extensive fine-tuning. A concept that is particularly advantageous in the medical imaging domain, where large annotated datasets for traditional training are scarce. In this context, a comprehensive evaluation of foundation models for polyp segmentation was conducted, assessing both detection and delimitation. For the study, three different colonoscopy datasets have been employed to compare the performance of five different foundation models, DINOv2, YOLO-World, GroundingDINO, SAM and MedSAM, against two benchmark networks, YOLOv8 and Mask R-CNN. Results show that the success of foundation models in polyp characterization is highly dependent on domain specialization. For optimal performance in medical applications, domain-specific models are essential, and generic models require fine-tuning to achieve effective results. Through this specialization, foundation models demonstrated superior performance compared to state-of-the-art detection and segmentation models, with some models even excelling in zero-shot evaluation; outperforming fine-tuned models on unseen data.
format Preprint
id arxiv_https___arxiv_org_abs_2503_24138
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AI-Assisted Colonoscopy: Polyp Detection and Segmentation using Foundation Models
Delaquintana-Aramendi, Uxue
Benito-del-Valle, Leire
Alvarez-Gila, Aitor
Pascau, Javier
Sánchez-Peralta, Luisa F
Picón, Artzai
Pagador, J Blas
Saratxaga, Cristina L
Image and Video Processing
Computer Vision and Pattern Recognition
In colonoscopy, 80% of the missed polyps could be detected with the help of Deep Learning models. In the search for algorithms capable of addressing this challenge, foundation models emerge as promising candidates. Their zero-shot or few-shot learning capabilities, facilitate generalization to new data or tasks without extensive fine-tuning. A concept that is particularly advantageous in the medical imaging domain, where large annotated datasets for traditional training are scarce. In this context, a comprehensive evaluation of foundation models for polyp segmentation was conducted, assessing both detection and delimitation. For the study, three different colonoscopy datasets have been employed to compare the performance of five different foundation models, DINOv2, YOLO-World, GroundingDINO, SAM and MedSAM, against two benchmark networks, YOLOv8 and Mask R-CNN. Results show that the success of foundation models in polyp characterization is highly dependent on domain specialization. For optimal performance in medical applications, domain-specific models are essential, and generic models require fine-tuning to achieve effective results. Through this specialization, foundation models demonstrated superior performance compared to state-of-the-art detection and segmentation models, with some models even excelling in zero-shot evaluation; outperforming fine-tuned models on unseen data.
title AI-Assisted Colonoscopy: Polyp Detection and Segmentation using Foundation Models
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.24138