ZIM: Zero-Shot Image Matting for Anything

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Beomyoung, Shin, Chanyong, Jeong, Joonhyun, Jung, Hyungsik, Lee, Se-Yun, Chun, Sewhan, Hwang, Dong-Hyun, Yu, Joonsang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911125947285504
author Kim, Beomyoung
Shin, Chanyong
Jeong, Joonhyun
Jung, Hyungsik
Lee, Se-Yun
Chun, Sewhan
Hwang, Dong-Hyun
Yu, Joonsang
author_facet Kim, Beomyoung
Shin, Chanyong
Jeong, Joonhyun
Jung, Hyungsik
Lee, Se-Yun
Chun, Sewhan
Hwang, Dong-Hyun
Yu, Joonsang
contents The recent segmentation foundation model, Segment Anything Model (SAM), exhibits strong zero-shot segmentation capabilities, but it falls short in generating fine-grained precise masks. To address this limitation, we propose a novel zero-shot image matting model, called ZIM, with two key contributions: First, we develop a label converter that transforms segmentation labels into detailed matte labels, constructing the new SA1B-Matte dataset without costly manual annotations. Training SAM with this dataset enables it to generate precise matte masks while maintaining its zero-shot capability. Second, we design the zero-shot matting model equipped with a hierarchical pixel decoder to enhance mask representation, along with a prompt-aware masked attention mechanism to improve performance by enabling the model to focus on regions specified by visual prompts. We evaluate ZIM using the newly introduced MicroMat-3K test set, which contains high-quality micro-level matte labels. Experimental results show that ZIM outperforms existing methods in fine-grained mask generation and zero-shot generalization. Furthermore, we demonstrate the versatility of ZIM in various downstream tasks requiring precise masks, such as image inpainting and 3D NeRF. Our contributions provide a robust foundation for advancing zero-shot matting and its downstream applications across a wide range of computer vision tasks. The code is available at https://github.com/naver-ai/ZIM.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00626
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ZIM: Zero-Shot Image Matting for Anything
Kim, Beomyoung
Shin, Chanyong
Jeong, Joonhyun
Jung, Hyungsik
Lee, Se-Yun
Chun, Sewhan
Hwang, Dong-Hyun
Yu, Joonsang
Computer Vision and Pattern Recognition
The recent segmentation foundation model, Segment Anything Model (SAM), exhibits strong zero-shot segmentation capabilities, but it falls short in generating fine-grained precise masks. To address this limitation, we propose a novel zero-shot image matting model, called ZIM, with two key contributions: First, we develop a label converter that transforms segmentation labels into detailed matte labels, constructing the new SA1B-Matte dataset without costly manual annotations. Training SAM with this dataset enables it to generate precise matte masks while maintaining its zero-shot capability. Second, we design the zero-shot matting model equipped with a hierarchical pixel decoder to enhance mask representation, along with a prompt-aware masked attention mechanism to improve performance by enabling the model to focus on regions specified by visual prompts. We evaluate ZIM using the newly introduced MicroMat-3K test set, which contains high-quality micro-level matte labels. Experimental results show that ZIM outperforms existing methods in fine-grained mask generation and zero-shot generalization. Furthermore, we demonstrate the versatility of ZIM in various downstream tasks requiring precise masks, such as image inpainting and 3D NeRF. Our contributions provide a robust foundation for advancing zero-shot matting and its downstream applications across a wide range of computer vision tasks. The code is available at https://github.com/naver-ai/ZIM.
title ZIM: Zero-Shot Image Matting for Anything
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.00626