Mixed-Query Transformer: A Unified Image Segmentation Architecture

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Pei, Cai, Zhaowei, Yang, Hao, Swaminathan, Ashwin, Manmatha, R., Soatto, Stefano
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913302405185536
author Wang, Pei
Cai, Zhaowei
Yang, Hao
Swaminathan, Ashwin
Manmatha, R.
Soatto, Stefano
author_facet Wang, Pei
Cai, Zhaowei
Yang, Hao
Swaminathan, Ashwin
Manmatha, R.
Soatto, Stefano
contents Existing unified image segmentation models either employ a unified architecture across multiple tasks but use separate weights tailored to each dataset, or apply a single set of weights to multiple datasets but are limited to a single task. In this paper, we introduce the Mixed-Query Transformer (MQ-Former), a unified architecture for multi-task and multi-dataset image segmentation using a single set of weights. To enable this, we propose a mixed query strategy, which can effectively and dynamically accommodate different types of objects without heuristic designs. In addition, the unified architecture allows us to use data augmentation with synthetic masks and captions to further improve model generalization. Experiments demonstrate that MQ-Former can not only effectively handle multiple segmentation datasets and tasks compared to specialized state-of-the-art models with competitive performance, but also generalize better to open-set segmentation tasks, evidenced by over 7 points higher performance than the prior art on the open-vocabulary SeginW benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04469
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mixed-Query Transformer: A Unified Image Segmentation Architecture
Wang, Pei
Cai, Zhaowei
Yang, Hao
Swaminathan, Ashwin
Manmatha, R.
Soatto, Stefano
Computer Vision and Pattern Recognition
Existing unified image segmentation models either employ a unified architecture across multiple tasks but use separate weights tailored to each dataset, or apply a single set of weights to multiple datasets but are limited to a single task. In this paper, we introduce the Mixed-Query Transformer (MQ-Former), a unified architecture for multi-task and multi-dataset image segmentation using a single set of weights. To enable this, we propose a mixed query strategy, which can effectively and dynamically accommodate different types of objects without heuristic designs. In addition, the unified architecture allows us to use data augmentation with synthetic masks and captions to further improve model generalization. Experiments demonstrate that MQ-Former can not only effectively handle multiple segmentation datasets and tasks compared to specialized state-of-the-art models with competitive performance, but also generalize better to open-set segmentation tasks, evidenced by over 7 points higher performance than the prior art on the open-vocabulary SeginW benchmark.
title Mixed-Query Transformer: A Unified Image Segmentation Architecture
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.04469