Your ViT is Secretly an Image Segmentation Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kerssies, Tommie, Cavagnero, Niccolò, Hermans, Alexander, Norouzi, Narges, Averta, Giuseppe, Leibe, Bastian, Dubbelman, Gijs, de Geus, Daan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913756648308736
author Kerssies, Tommie
Cavagnero, Niccolò
Hermans, Alexander
Norouzi, Narges
Averta, Giuseppe
Leibe, Bastian
Dubbelman, Gijs
de Geus, Daan
author_facet Kerssies, Tommie
Cavagnero, Niccolò
Hermans, Alexander
Norouzi, Narges
Averta, Giuseppe
Leibe, Bastian
Dubbelman, Gijs
de Geus, Daan
contents Vision Transformers (ViTs) have shown remarkable performance and scalability across various computer vision tasks. To apply single-scale ViTs to image segmentation, existing methods adopt a convolutional adapter to generate multi-scale features, a pixel decoder to fuse these features, and a Transformer decoder that uses the fused features to make predictions. In this paper, we show that the inductive biases introduced by these task-specific components can instead be learned by the ViT itself, given sufficiently large models and extensive pre-training. Based on these findings, we introduce the Encoder-only Mask Transformer (EoMT), which repurposes the plain ViT architecture to conduct image segmentation. With large-scale models and pre-training, EoMT obtains a segmentation accuracy similar to state-of-the-art models that use task-specific components. At the same time, EoMT is significantly faster than these methods due to its architectural simplicity, e.g., up to 4x faster with ViT-L. Across a range of model sizes, EoMT demonstrates an optimal balance between segmentation accuracy and prediction speed, suggesting that compute resources are better spent on scaling the ViT itself rather than adding architectural complexity. Code: https://www.tue-mps.org/eomt/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19108
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Your ViT is Secretly an Image Segmentation Model
Kerssies, Tommie
Cavagnero, Niccolò
Hermans, Alexander
Norouzi, Narges
Averta, Giuseppe
Leibe, Bastian
Dubbelman, Gijs
de Geus, Daan
Computer Vision and Pattern Recognition
Vision Transformers (ViTs) have shown remarkable performance and scalability across various computer vision tasks. To apply single-scale ViTs to image segmentation, existing methods adopt a convolutional adapter to generate multi-scale features, a pixel decoder to fuse these features, and a Transformer decoder that uses the fused features to make predictions. In this paper, we show that the inductive biases introduced by these task-specific components can instead be learned by the ViT itself, given sufficiently large models and extensive pre-training. Based on these findings, we introduce the Encoder-only Mask Transformer (EoMT), which repurposes the plain ViT architecture to conduct image segmentation. With large-scale models and pre-training, EoMT obtains a segmentation accuracy similar to state-of-the-art models that use task-specific components. At the same time, EoMT is significantly faster than these methods due to its architectural simplicity, e.g., up to 4x faster with ViT-L. Across a range of model sizes, EoMT demonstrates an optimal balance between segmentation accuracy and prediction speed, suggesting that compute resources are better spent on scaling the ViT itself rather than adding architectural complexity. Code: https://www.tue-mps.org/eomt/.
title Your ViT is Secretly an Image Segmentation Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.19108