VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Lehan, Song, Jincen, Wang, Tianlong, Qi, Daiqing, Shi, Weili, Liu, Yuheng, Li, Sheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929759252905984
author Yang, Lehan
Song, Jincen
Wang, Tianlong
Qi, Daiqing
Shi, Weili
Liu, Yuheng
Li, Sheng
author_facet Yang, Lehan
Song, Jincen
Wang, Tianlong
Qi, Daiqing
Shi, Weili
Liu, Yuheng
Li, Sheng
contents We propose a new task, video referring matting, which obtains the alpha matte of a specified instance by inputting a referring caption. We treat the dense prediction task of matting as video generation, leveraging the text-to-video alignment prior of video diffusion models to generate alpha mattes that are temporally coherent and closely related to the corresponding semantic instances. Moreover, we propose a new Latent-Constructive loss to further distinguish different instances, enabling more controllable interactive matting. Additionally, we introduce a large-scale video referring matting dataset with 10,000 videos. To the best of our knowledge, this is the first dataset that concurrently contains captions, videos, and instance-level alpha mattes. Extensive experiments demonstrate the effectiveness of our method. The dataset and code are available at https://github.com/Hansxsourse/VRMDiff.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10678
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
Yang, Lehan
Song, Jincen
Wang, Tianlong
Qi, Daiqing
Shi, Weili
Liu, Yuheng
Li, Sheng
Computer Vision and Pattern Recognition
We propose a new task, video referring matting, which obtains the alpha matte of a specified instance by inputting a referring caption. We treat the dense prediction task of matting as video generation, leveraging the text-to-video alignment prior of video diffusion models to generate alpha mattes that are temporally coherent and closely related to the corresponding semantic instances. Moreover, we propose a new Latent-Constructive loss to further distinguish different instances, enabling more controllable interactive matting. Additionally, we introduce a large-scale video referring matting dataset with 10,000 videos. To the best of our knowledge, this is the first dataset that concurrently contains captions, videos, and instance-level alpha mattes. Extensive experiments demonstrate the effectiveness of our method. The dataset and code are available at https://github.com/Hansxsourse/VRMDiff.
title VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.10678