PixelLM: Pixel Reasoning with Large Multimodal Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Zhongwei, Huang, Zhicheng, Wei, Yunchao, Zhao, Yao, Fu, Dongmei, Feng, Jiashi, Jin, Xiaojie
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916328992931840
author Ren, Zhongwei
Huang, Zhicheng
Wei, Yunchao
Zhao, Yao
Fu, Dongmei
Feng, Jiashi
Jin, Xiaojie
author_facet Ren, Zhongwei
Huang, Zhicheng
Wei, Yunchao
Zhao, Yao
Fu, Dongmei
Feng, Jiashi
Jin, Xiaojie
contents While large multimodal models (LMMs) have achieved remarkable progress, generating pixel-level masks for image reasoning tasks involving multiple open-world targets remains a challenge. To bridge this gap, we introduce PixelLM, an effective and efficient LMM for pixel-level reasoning and understanding. Central to PixelLM is a novel, lightweight pixel decoder and a comprehensive segmentation codebook. The decoder efficiently produces masks from the hidden embeddings of the codebook tokens, which encode detailed target-relevant information. With this design, PixelLM harmonizes with the structure of popular LMMs and avoids the need for additional costly segmentation models. Furthermore, we propose a target refinement loss to enhance the model's ability to differentiate between multiple targets, leading to substantially improved mask quality. To advance research in this area, we construct MUSE, a high-quality multi-target reasoning segmentation benchmark. PixelLM excels across various pixel-level image reasoning and understanding tasks, outperforming well-established methods in multiple benchmarks, including MUSE, single- and multi-referring segmentation. Comprehensive ablations confirm the efficacy of each proposed component. All code, models, and datasets will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2312_02228
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle PixelLM: Pixel Reasoning with Large Multimodal Model
Ren, Zhongwei
Huang, Zhicheng
Wei, Yunchao
Zhao, Yao
Fu, Dongmei
Feng, Jiashi
Jin, Xiaojie
Computer Vision and Pattern Recognition
While large multimodal models (LMMs) have achieved remarkable progress, generating pixel-level masks for image reasoning tasks involving multiple open-world targets remains a challenge. To bridge this gap, we introduce PixelLM, an effective and efficient LMM for pixel-level reasoning and understanding. Central to PixelLM is a novel, lightweight pixel decoder and a comprehensive segmentation codebook. The decoder efficiently produces masks from the hidden embeddings of the codebook tokens, which encode detailed target-relevant information. With this design, PixelLM harmonizes with the structure of popular LMMs and avoids the need for additional costly segmentation models. Furthermore, we propose a target refinement loss to enhance the model's ability to differentiate between multiple targets, leading to substantially improved mask quality. To advance research in this area, we construct MUSE, a high-quality multi-target reasoning segmentation benchmark. PixelLM excels across various pixel-level image reasoning and understanding tasks, outperforming well-established methods in multiple benchmarks, including MUSE, single- and multi-referring segmentation. Comprehensive ablations confirm the efficacy of each proposed component. All code, models, and datasets will be publicly available.
title PixelLM: Pixel Reasoning with Large Multimodal Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.02228