Quickly Tuning Foundation Models for Image Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Das, Breenda, Purucker, Lennart, Carstensen, Timur, Hutter, Frank
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916914301763584
author Das, Breenda
Purucker, Lennart
Carstensen, Timur
Hutter, Frank
author_facet Das, Breenda
Purucker, Lennart
Carstensen, Timur
Hutter, Frank
contents Foundation models like SAM (Segment Anything Model) exhibit strong zero-shot image segmentation performance, but often fall short on domain-specific tasks. Fine-tuning these models typically requires significant manual effort and domain expertise. In this work, we introduce QTT-SEG, a meta-learning-driven approach for automating and accelerating the fine-tuning of SAM for image segmentation. Built on the Quick-Tune hyperparameter optimization framework, QTT-SEG predicts high-performing configurations using meta-learned cost and performance models, efficiently navigating a search space of over 200 million possibilities. We evaluate QTT-SEG on eight binary and five multiclass segmentation datasets under tight time constraints. Our results show that QTT-SEG consistently improves upon SAM's zero-shot performance and surpasses AutoGluon Multimodal, a strong AutoML baseline, on most binary tasks within three minutes. On multiclass datasets, QTT-SEG delivers consistent gains as well. These findings highlight the promise of meta-learning in automating model adaptation for specialized segmentation tasks. Code available at: https://github.com/ds-brx/QTT-SEG/
format Preprint
id arxiv_https___arxiv_org_abs_2508_17283
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quickly Tuning Foundation Models for Image Segmentation
Das, Breenda
Purucker, Lennart
Carstensen, Timur
Hutter, Frank
Computer Vision and Pattern Recognition
Machine Learning
Foundation models like SAM (Segment Anything Model) exhibit strong zero-shot image segmentation performance, but often fall short on domain-specific tasks. Fine-tuning these models typically requires significant manual effort and domain expertise. In this work, we introduce QTT-SEG, a meta-learning-driven approach for automating and accelerating the fine-tuning of SAM for image segmentation. Built on the Quick-Tune hyperparameter optimization framework, QTT-SEG predicts high-performing configurations using meta-learned cost and performance models, efficiently navigating a search space of over 200 million possibilities. We evaluate QTT-SEG on eight binary and five multiclass segmentation datasets under tight time constraints. Our results show that QTT-SEG consistently improves upon SAM's zero-shot performance and surpasses AutoGluon Multimodal, a strong AutoML baseline, on most binary tasks within three minutes. On multiclass datasets, QTT-SEG delivers consistent gains as well. These findings highlight the promise of meta-learning in automating model adaptation for specialized segmentation tasks. Code available at: https://github.com/ds-brx/QTT-SEG/
title Quickly Tuning Foundation Models for Image Segmentation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2508.17283