Text4Seg: Reimagining Image Segmentation as Text Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lan, Mengcheng, Chen, Chaofeng, Zhou, Yue, Xu, Jiaxing, Ke, Yiping, Wang, Xinjiang, Feng, Litong, Zhang, Wayne
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915153597956096
author Lan, Mengcheng
Chen, Chaofeng
Zhou, Yue
Xu, Jiaxing
Ke, Yiping
Wang, Xinjiang
Feng, Litong
Zhang, Wayne
author_facet Lan, Mengcheng
Chen, Chaofeng
Zhou, Yue
Xu, Jiaxing
Ke, Yiping
Wang, Xinjiang
Feng, Litong
Zhang, Wayne
contents Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains a significant challenge. In this paper, we introduce Text4Seg, a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. This unified representation allows seamless integration into the auto-regressive training pipeline of MLLMs for easier optimization. We demonstrate that representing an image with $16\times16$ semantic descriptors yields competitive segmentation performance. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptors by 74% and accelerating inference by $3\times$, without compromising performance. Extensive experiments across various vision tasks, such as referring expression segmentation and comprehension, show that Text4Seg achieves state-of-the-art performance on multiple datasets by fine-tuning different MLLM backbones. Our approach provides an efficient, scalable solution for vision-centric tasks within the MLLM framework.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09855
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text4Seg: Reimagining Image Segmentation as Text Generation
Lan, Mengcheng
Chen, Chaofeng
Zhou, Yue
Xu, Jiaxing
Ke, Yiping
Wang, Xinjiang
Feng, Litong
Zhang, Wayne
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains a significant challenge. In this paper, we introduce Text4Seg, a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. This unified representation allows seamless integration into the auto-regressive training pipeline of MLLMs for easier optimization. We demonstrate that representing an image with $16\times16$ semantic descriptors yields competitive segmentation performance. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptors by 74% and accelerating inference by $3\times$, without compromising performance. Extensive experiments across various vision tasks, such as referring expression segmentation and comprehension, show that Text4Seg achieves state-of-the-art performance on multiple datasets by fine-tuning different MLLM backbones. Our approach provides an efficient, scalable solution for vision-centric tasks within the MLLM framework.
title Text4Seg: Reimagining Image Segmentation as Text Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.09855