Saved in:
Bibliographic Details
Main Authors: Zhang, Zheng, Ma, Yeyao, Zhang, Enming, Bai, Xiang
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.14598
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909145155764224
author Zhang, Zheng
Ma, Yeyao
Zhang, Enming
Bai, Xiang
author_facet Zhang, Zheng
Ma, Yeyao
Zhang, Enming
Bai, Xiang
contents PSALM is a powerful extension of the Large Multi-modal Model (LMM) to address the segmentation task challenges. To overcome the limitation of the LMM being limited to textual output, PSALM incorporates a mask decoder and a well-designed input schema to handle a variety of segmentation tasks. This schema includes images, task instructions, conditional prompts, and mask tokens, which enable the model to generate and classify segmentation masks effectively. The flexible design of PSALM supports joint training across multiple datasets and tasks, leading to improved performance and task generalization. PSALM achieves superior results on several benchmarks, such as RefCOCO/RefCOCO+/RefCOCOg, COCO Panoptic Segmentation, and COCO-Interactive, and further exhibits zero-shot capabilities on unseen tasks, such as open-vocabulary segmentation, generalized referring expression segmentation and video object segmentation, making a significant step towards a GPT moment in computer vision. Through extensive experiments, PSALM demonstrates its potential to transform the domain of image segmentation, leveraging the robust visual understanding capabilities of LMMs as seen in natural language processing. Code and models are available at https://github.com/zamling/PSALM.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14598
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model
Zhang, Zheng
Ma, Yeyao
Zhang, Enming
Bai, Xiang
Computer Vision and Pattern Recognition
PSALM is a powerful extension of the Large Multi-modal Model (LMM) to address the segmentation task challenges. To overcome the limitation of the LMM being limited to textual output, PSALM incorporates a mask decoder and a well-designed input schema to handle a variety of segmentation tasks. This schema includes images, task instructions, conditional prompts, and mask tokens, which enable the model to generate and classify segmentation masks effectively. The flexible design of PSALM supports joint training across multiple datasets and tasks, leading to improved performance and task generalization. PSALM achieves superior results on several benchmarks, such as RefCOCO/RefCOCO+/RefCOCOg, COCO Panoptic Segmentation, and COCO-Interactive, and further exhibits zero-shot capabilities on unseen tasks, such as open-vocabulary segmentation, generalized referring expression segmentation and video object segmentation, making a significant step towards a GPT moment in computer vision. Through extensive experiments, PSALM demonstrates its potential to transform the domain of image segmentation, leveraging the robust visual understanding capabilities of LMMs as seen in natural language processing. Code and models are available at https://github.com/zamling/PSALM.
title PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.14598