VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929759252905984 |
|---|---|
| author | Yang, Lehan Song, Jincen Wang, Tianlong Qi, Daiqing Shi, Weili Liu, Yuheng Li, Sheng |
| author_facet | Yang, Lehan Song, Jincen Wang, Tianlong Qi, Daiqing Shi, Weili Liu, Yuheng Li, Sheng |
| contents | We propose a new task, video referring matting, which obtains the alpha matte of a specified instance by inputting a referring caption. We treat the dense prediction task of matting as video generation, leveraging the text-to-video alignment prior of video diffusion models to generate alpha mattes that are temporally coherent and closely related to the corresponding semantic instances. Moreover, we propose a new Latent-Constructive loss to further distinguish different instances, enabling more controllable interactive matting. Additionally, we introduce a large-scale video referring matting dataset with 10,000 videos. To the best of our knowledge, this is the first dataset that concurrently contains captions, videos, and instance-level alpha mattes. Extensive experiments demonstrate the effectiveness of our method. The dataset and code are available at https://github.com/Hansxsourse/VRMDiff. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_10678 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion Yang, Lehan Song, Jincen Wang, Tianlong Qi, Daiqing Shi, Weili Liu, Yuheng Li, Sheng Computer Vision and Pattern Recognition We propose a new task, video referring matting, which obtains the alpha matte of a specified instance by inputting a referring caption. We treat the dense prediction task of matting as video generation, leveraging the text-to-video alignment prior of video diffusion models to generate alpha mattes that are temporally coherent and closely related to the corresponding semantic instances. Moreover, we propose a new Latent-Constructive loss to further distinguish different instances, enabling more controllable interactive matting. Additionally, we introduce a large-scale video referring matting dataset with 10,000 videos. To the best of our knowledge, this is the first dataset that concurrently contains captions, videos, and instance-level alpha mattes. Extensive experiments demonstrate the effectiveness of our method. The dataset and code are available at https://github.com/Hansxsourse/VRMDiff. |
| title | VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.10678 |