Saved in:
Bibliographic Details
Main Authors: Huang, Longfei, Liang, Yu, Zhang, Hao, Chen, Jinwei, Dong, Wei, Chen, Lunde, Liu, Wanyu, Li, Bo, Jiang, Peng-Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.00443
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913974062153728
author Huang, Longfei
Liang, Yu
Zhang, Hao
Chen, Jinwei
Dong, Wei
Chen, Lunde
Liu, Wanyu
Li, Bo
Jiang, Peng-Tao
author_facet Huang, Longfei
Liang, Yu
Zhang, Hao
Chen, Jinwei
Dong, Wei
Chen, Lunde
Liu, Wanyu
Li, Bo
Jiang, Peng-Tao
contents Recent interactive matting methods have shown satisfactory performance in capturing the primary regions of objects, but they fall short in extracting fine-grained details in edge regions. Diffusion models trained on billions of image-text pairs, demonstrate exceptional capability in modeling highly complex data distributions and synthesizing realistic texture details, while exhibiting robust text-driven interaction capabilities, making them an attractive solution for interactive matting. To this end, we propose SDMatte, a diffusion-driven interactive matting model, with three key contributions. First, we exploit the powerful priors of diffusion models and transform the text-driven interaction capability into visual prompt-driven interaction capability to enable interactive matting. Second, we integrate coordinate embeddings of visual prompts and opacity embeddings of target objects into U-Net, enhancing SDMatte's sensitivity to spatial position information and opacity information. Third, we propose a masked self-attention mechanism that enables the model to focus on areas specified by visual prompts, leading to better performance. Extensive experiments on multiple datasets demonstrate the superior performance of our method, validating its effectiveness in interactive matting. Our code and model are available at https://github.com/vivoCameraResearch/SDMatte.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00443
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SDMatte: Grafting Diffusion Models for Interactive Matting
Huang, Longfei
Liang, Yu
Zhang, Hao
Chen, Jinwei
Dong, Wei
Chen, Lunde
Liu, Wanyu
Li, Bo
Jiang, Peng-Tao
Computer Vision and Pattern Recognition
Recent interactive matting methods have shown satisfactory performance in capturing the primary regions of objects, but they fall short in extracting fine-grained details in edge regions. Diffusion models trained on billions of image-text pairs, demonstrate exceptional capability in modeling highly complex data distributions and synthesizing realistic texture details, while exhibiting robust text-driven interaction capabilities, making them an attractive solution for interactive matting. To this end, we propose SDMatte, a diffusion-driven interactive matting model, with three key contributions. First, we exploit the powerful priors of diffusion models and transform the text-driven interaction capability into visual prompt-driven interaction capability to enable interactive matting. Second, we integrate coordinate embeddings of visual prompts and opacity embeddings of target objects into U-Net, enhancing SDMatte's sensitivity to spatial position information and opacity information. Third, we propose a masked self-attention mechanism that enables the model to focus on areas specified by visual prompts, leading to better performance. Extensive experiments on multiple datasets demonstrate the superior performance of our method, validating its effectiveness in interactive matting. Our code and model are available at https://github.com/vivoCameraResearch/SDMatte.
title SDMatte: Grafting Diffusion Models for Interactive Matting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.00443