VLSM-Ensemble: Ensembling CLIP-based Vision-Language Models for Enhanced Medical Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dietlmeier, Julia, Adegboro, Oluwabukola Grace, Ganepola, Vayangi, Mazo, Claudia, O'Connor, Noel E.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916935930740736
author Dietlmeier, Julia
Adegboro, Oluwabukola Grace
Ganepola, Vayangi
Mazo, Claudia
O'Connor, Noel E.
author_facet Dietlmeier, Julia
Adegboro, Oluwabukola Grace
Ganepola, Vayangi
Mazo, Claudia
O'Connor, Noel E.
contents Vision-language models and their adaptations to image segmentation tasks present enormous potential for producing highly accurate and interpretable results. However, implementations based on CLIP and BiomedCLIP are still lagging behind more sophisticated architectures such as CRIS. In this work, instead of focusing on text prompt engineering as is the norm, we attempt to narrow this gap by showing how to ensemble vision-language segmentation models (VLSMs) with a low-complexity CNN. By doing so, we achieve a significant Dice score improvement of 6.3% on the BKAI polyp dataset using the ensembled BiomedCLIPSeg, while other datasets exhibit gains ranging from 1% to 6%. Furthermore, we provide initial results on additional four radiology and non-radiology datasets. We conclude that ensembling works differently across these datasets (from outperforming to underperforming the CRIS model), indicating a topic for future investigation by the community. The code is available at https://github.com/juliadietlmeier/VLSM-Ensemble.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05154
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLSM-Ensemble: Ensembling CLIP-based Vision-Language Models for Enhanced Medical Image Segmentation
Dietlmeier, Julia
Adegboro, Oluwabukola Grace
Ganepola, Vayangi
Mazo, Claudia
O'Connor, Noel E.
Image and Video Processing
Computer Vision and Pattern Recognition
Vision-language models and their adaptations to image segmentation tasks present enormous potential for producing highly accurate and interpretable results. However, implementations based on CLIP and BiomedCLIP are still lagging behind more sophisticated architectures such as CRIS. In this work, instead of focusing on text prompt engineering as is the norm, we attempt to narrow this gap by showing how to ensemble vision-language segmentation models (VLSMs) with a low-complexity CNN. By doing so, we achieve a significant Dice score improvement of 6.3% on the BKAI polyp dataset using the ensembled BiomedCLIPSeg, while other datasets exhibit gains ranging from 1% to 6%. Furthermore, we provide initial results on additional four radiology and non-radiology datasets. We conclude that ensembling works differently across these datasets (from outperforming to underperforming the CRIS model), indicating a topic for future investigation by the community. The code is available at https://github.com/juliadietlmeier/VLSM-Ensemble.
title VLSM-Ensemble: Ensembling CLIP-based Vision-Language Models for Enhanced Medical Image Segmentation
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.05154