VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Lehan, Song, Jincen, Wang, Tianlong, Qi, Daiqing, Shi, Weili, Liu, Yuheng, Li, Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
by: Huang, Binyuan, et al.
Published: (2026)
by: Huang, Binyuan, et al.
Published: (2026)
Generative Video Matting
by: Ge, Yongtao, et al.
Published: (2025)
by: Ge, Yongtao, et al.
Published: (2025)
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
Diffusion for Natural Image Matting
by: Hu, Yihan, et al.
Published: (2023)
by: Hu, Yihan, et al.
Published: (2023)
Matting by Generation
by: Wang, Zhixiang, et al.
Published: (2024)
by: Wang, Zhixiang, et al.
Published: (2024)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
by: Feng, Weixi, et al.
Published: (2025)
by: Feng, Weixi, et al.
Published: (2025)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
by: Qi, Daiqing, et al.
Published: (2024)
by: Qi, Daiqing, et al.
Published: (2024)
Dual-Stream Diffusion Net for Text-to-Video Generation
by: Liu, Binhui, et al.
Published: (2023)
by: Liu, Binhui, et al.
Published: (2023)
MatAnyone: Stable Video Matting with Consistent Memory Propagation
by: Yang, Peiqing, et al.
Published: (2025)
by: Yang, Peiqing, et al.
Published: (2025)
VideoMaMa: Mask-Guided Video Matting via Generative Prior
by: Lim, Sangbeom, et al.
Published: (2026)
by: Lim, Sangbeom, et al.
Published: (2026)
VideoMat: Extracting PBR Materials from Video Diffusion Models
by: Munkberg, Jacob, et al.
Published: (2025)
by: Munkberg, Jacob, et al.
Published: (2025)
GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors
by: Yin, Xingyilang, et al.
Published: (2025)
by: Yin, Xingyilang, et al.
Published: (2025)
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
by: Wang, Wenjing, et al.
Published: (2023)
by: Wang, Wenjing, et al.
Published: (2023)
Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
by: Wu, Daiqing, et al.
Published: (2025)
by: Wu, Daiqing, et al.
Published: (2025)
Pyramid Diffusion for Fine 3D Large Scene Generation
by: Liu, Yuheng, et al.
Published: (2023)
by: Liu, Yuheng, et al.
Published: (2023)
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
by: Zhang, Yan, et al.
Published: (2024)
by: Zhang, Yan, et al.
Published: (2024)
OptiWorld: Optimal Control for Video World Generation under Physical Constraints
by: Yuan, Yu, et al.
Published: (2026)
by: Yuan, Yu, et al.
Published: (2026)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
by: Han, Wenkang, et al.
Published: (2025)
by: Han, Wenkang, et al.
Published: (2025)
TextDiff: Mask-Guided Residual Diffusion Models for Scene Text Image Super-Resolution
by: Liu, Baolin, et al.
Published: (2023)
by: Liu, Baolin, et al.
Published: (2023)
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
by: Zhang, Ruixin, et al.
Published: (2025)
by: Zhang, Ruixin, et al.
Published: (2025)
MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator
by: Yang, Peiqing, et al.
Published: (2025)
by: Yang, Peiqing, et al.
Published: (2025)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
PixelDiT: Pixel Diffusion Transformers for Image Generation
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Conditional Text-to-Image Generation with Reference Guidance
by: Kim, Taewook, et al.
Published: (2024)
by: Kim, Taewook, et al.
Published: (2024)
VIRES: Video Instance Repainting via Sketch and Text Guided Generation
by: Weng, Shuchen, et al.
Published: (2024)
by: Weng, Shuchen, et al.
Published: (2024)
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
by: Ren, Yixuan, et al.
Published: (2024)
by: Ren, Yixuan, et al.
Published: (2024)
Weakly-Supervised Referring Video Object Segmentation through Text Supervision
by: Shi, Miaojing, et al.
Published: (2026)
by: Shi, Miaojing, et al.
Published: (2026)
Beyond Detection: A Structure-Aware Framework for Scene Text Tracking
by: Yu, Chenmin, et al.
Published: (2026)
by: Yu, Chenmin, et al.
Published: (2026)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
by: Zhu, Zixin, et al.
Published: (2024)
by: Zhu, Zixin, et al.
Published: (2024)
Text-Animator: Controllable Visual Text Video Generation
by: Liu, Lin, et al.
Published: (2024)
by: Liu, Lin, et al.
Published: (2024)
MatPhys: Learning Material-Aware Physics Parameters for Deformable Object Simulation from Videos
by: Yang, Yang, et al.
Published: (2026)
by: Yang, Yang, et al.
Published: (2026)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
SDMatte: Grafting Diffusion Models for Interactive Matting
by: Huang, Longfei, et al.
Published: (2025)
by: Huang, Longfei, et al.
Published: (2025)
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
by: Dou, Huanzhang, et al.
Published: (2024)
by: Dou, Huanzhang, et al.
Published: (2024)
RobustMat: Neural Diffusion for Street Landmark Patch Matching under Challenging Environments
by: She, Rui, et al.
Published: (2023)
by: She, Rui, et al.
Published: (2023)
RefAlign: Representation Alignment for Reference-to-Video Generation
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Elevating Flow-Guided Video Inpainting with Reference Generation
by: Cho, Suhwan, et al.
Published: (2024)
by: Cho, Suhwan, et al.
Published: (2024)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
by: Peng, Bo, et al.
Published: (2023)
by: Peng, Bo, et al.
Published: (2023)
Similar Items
-
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025) -
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
by: Huang, Binyuan, et al.
Published: (2026) -
Generative Video Matting
by: Ge, Yongtao, et al.
Published: (2025) -
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
by: Chen, Yuheng, et al.
Published: (2026) -
Diffusion for Natural Image Matting
by: Hu, Yihan, et al.
Published: (2023)