Test-Time Optimization for Domain Adaptive Open Vocabulary Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: De Silva, Ulindu, Samaraweera, Didula, Wanigathunga, Sasini, Kariyawasam, Kavindu, Ranasinghe, Kanchana, Naseer, Muzammal, Rodrigo, Ranga
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912265349890048
author De Silva, Ulindu
Samaraweera, Didula
Wanigathunga, Sasini
Kariyawasam, Kavindu
Ranasinghe, Kanchana
Naseer, Muzammal
Rodrigo, Ranga
author_facet De Silva, Ulindu
Samaraweera, Didula
Wanigathunga, Sasini
Kariyawasam, Kavindu
Ranasinghe, Kanchana
Naseer, Muzammal
Rodrigo, Ranga
contents We present Seg-TTO, a novel framework for zero-shot, open-vocabulary semantic segmentation (OVSS), designed to excel in specialized domain tasks. While current open-vocabulary approaches show impressive performance on standard segmentation benchmarks under zero-shot settings, they fall short of supervised counterparts on highly domain-specific datasets. We focus on segmentation-specific test-time optimization to address this gap. Segmentation requires an understanding of multiple concepts within a single image while retaining the locality and spatial structure of representations. We propose a novel self-supervised objective adhering to these requirements and use it to align the model parameters with input images at test time. In the textual modality, we learn multiple embeddings for each category to capture diverse concepts within an image, while in the visual modality, we calculate pixel-level losses followed by embedding aggregation operations specific to preserving spatial structure. Our resulting framework termed Seg-TTO is a plug-and-play module. We integrate Seg-TTO with three state-of-the-art OVSS approaches and evaluate across 22 challenging OVSS tasks covering a range of specialized domains. Our Seg-TTO demonstrates clear performance improvements (up to 27% mIoU increase on some datasets) establishing new state-of-the-art. Our code and models will be released publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04696
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Test-Time Optimization for Domain Adaptive Open Vocabulary Segmentation
De Silva, Ulindu
Samaraweera, Didula
Wanigathunga, Sasini
Kariyawasam, Kavindu
Ranasinghe, Kanchana
Naseer, Muzammal
Rodrigo, Ranga
Computer Vision and Pattern Recognition
We present Seg-TTO, a novel framework for zero-shot, open-vocabulary semantic segmentation (OVSS), designed to excel in specialized domain tasks. While current open-vocabulary approaches show impressive performance on standard segmentation benchmarks under zero-shot settings, they fall short of supervised counterparts on highly domain-specific datasets. We focus on segmentation-specific test-time optimization to address this gap. Segmentation requires an understanding of multiple concepts within a single image while retaining the locality and spatial structure of representations. We propose a novel self-supervised objective adhering to these requirements and use it to align the model parameters with input images at test time. In the textual modality, we learn multiple embeddings for each category to capture diverse concepts within an image, while in the visual modality, we calculate pixel-level losses followed by embedding aggregation operations specific to preserving spatial structure. Our resulting framework termed Seg-TTO is a plug-and-play module. We integrate Seg-TTO with three state-of-the-art OVSS approaches and evaluate across 22 challenging OVSS tasks covering a range of specialized domains. Our Seg-TTO demonstrates clear performance improvements (up to 27% mIoU increase on some datasets) establishing new state-of-the-art. Our code and models will be released publicly.
title Test-Time Optimization for Domain Adaptive Open Vocabulary Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04696