Rethinking Video Segmentation with Masked Video Consistency: Did the Model Learn as Intended?
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Chen, Guo, Qiang, Qu, Xiaochao, Liu, Luoqi, Liu, Ting |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic Segmentation on VSPW Dataset through Masked Video Consistency
by: Liang, Chen, et al.
Published: (2024)
by: Liang, Chen, et al.
Published: (2024)
MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
by: Huang, Jun, et al.
Published: (2025)
by: Huang, Jun, et al.
Published: (2025)
Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation
by: Lou, Zijie, et al.
Published: (2026)
by: Lou, Zijie, et al.
Published: (2026)
SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything Model
by: Yu, Chongkai, et al.
Published: (2024)
by: Yu, Chongkai, et al.
Published: (2024)
GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching
by: Liu, Minghao, et al.
Published: (2024)
by: Liu, Minghao, et al.
Published: (2024)
2nd Place Solution for MOSE Track in CVPR 2024 PVUW workshop: Complex Video Object Segmentation
by: Xu, Zhensong, et al.
Published: (2024)
by: Xu, Zhensong, et al.
Published: (2024)
3rd Place Solution for PVUW Challenge 2024: Video Panoptic Segmentation
by: Wu, Ruipu, et al.
Published: (2024)
by: Wu, Ruipu, et al.
Published: (2024)
MiVE: Multiscale Vision-language features for reference-guided video Editing
by: Wang, Tong, et al.
Published: (2026)
by: Wang, Tong, et al.
Published: (2026)
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
by: Lin, Yiheng, et al.
Published: (2025)
by: Lin, Yiheng, et al.
Published: (2025)
Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
by: Li, Hongxi, et al.
Published: (2026)
by: Li, Hongxi, et al.
Published: (2026)
FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation
by: Zhang, Zekang, et al.
Published: (2026)
by: Zhang, Zekang, et al.
Published: (2026)
Memory Efficient Matting with Adaptive Token Routing
by: Lin, Yiheng, et al.
Published: (2024)
by: Lin, Yiheng, et al.
Published: (2024)
DCEdit: Dual-Level Controlled Image Editing via Precisely Localized Semantics
by: Hu, Yihan, et al.
Published: (2025)
by: Hu, Yihan, et al.
Published: (2025)
PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation
by: Sha, Yuyang, et al.
Published: (2026)
by: Sha, Yuyang, et al.
Published: (2026)
TextMastero: Mastering High-Quality Scene Text Editing in Diverse Languages and Styles
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
On Exact Editing of Flow-Based Diffusion Models
by: Li, Zixiang, et al.
Published: (2025)
by: Li, Zixiang, et al.
Published: (2025)
Adversarially Masked Video Consistency for Unsupervised Domain Adaptation
by: Zhu, Xiaoyu, et al.
Published: (2024)
by: Zhu, Xiaoyu, et al.
Published: (2024)
MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
by: Ni, Jingcheng, et al.
Published: (2025)
by: Ni, Jingcheng, et al.
Published: (2025)
Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
by: Lin, Ci-Siang, et al.
Published: (2025)
by: Lin, Ci-Siang, et al.
Published: (2025)
Rethinking Video-Language Model from the Language Input Perspective
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
by: Zeng, Jianshu, et al.
Published: (2025)
by: Zeng, Jianshu, et al.
Published: (2025)
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
by: Wang, Zanyi, et al.
Published: (2025)
by: Wang, Zanyi, et al.
Published: (2025)
SMC++: Masked Learning of Unsupervised Video Semantic Compression
by: Tian, Yuan, et al.
Published: (2024)
by: Tian, Yuan, et al.
Published: (2024)
Fréchet Video Motion Distance: A Metric for Evaluating Motion Consistency in Videos
by: Liu, Jiahe, et al.
Published: (2024)
by: Liu, Jiahe, et al.
Published: (2024)
Exploring AIGC Video Quality: A Focus on Visual Harmony, Video-Text Consistency and Domain Distribution Gap
by: Qu, Bowen, et al.
Published: (2024)
by: Qu, Bowen, et al.
Published: (2024)
VideoMAC: Video Masked Autoencoders Meet ConvNets
by: Pei, Gensheng, et al.
Published: (2024)
by: Pei, Gensheng, et al.
Published: (2024)
1st Place Winner of the 2024 Pixel-level Video Understanding in the Wild (CVPR'24 PVUW) Challenge in Video Panoptic Segmentation and Best Long Video Consistency of Video Semantic Segmentation
by: Liu, Qingfeng, et al.
Published: (2024)
by: Liu, Qingfeng, et al.
Published: (2024)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
by: Wang, Hengkang, et al.
Published: (2025)
by: Wang, Hengkang, et al.
Published: (2025)
Static-Dynamic Class-level Perception Consistency in Video Semantic Segmentation
by: Cen, Zhigang, et al.
Published: (2024)
by: Cen, Zhigang, et al.
Published: (2024)
V-CORE: Temporally Consistent Video Understanding for Video-LLM
by: Kang, Zhengjian, et al.
Published: (2026)
by: Kang, Zhengjian, et al.
Published: (2026)
CRCL: Causal Representation Consistency Learning for Anomaly Detection in Surveillance Videos
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation
by: Choi, Sun-Hyuk, et al.
Published: (2025)
by: Choi, Sun-Hyuk, et al.
Published: (2025)
RepVideo: Rethinking Cross-Layer Representation for Video Generation
by: Si, Chenyang, et al.
Published: (2025)
by: Si, Chenyang, et al.
Published: (2025)
First-frame Supervised Video Polyp Segmentation via Propagative and Semantic Dual-teacher Network
by: Hu, Qiang, et al.
Published: (2024)
by: Hu, Qiang, et al.
Published: (2024)
Rethinking Efficient Crack Segmentation with Task-Aligned Structural-Directional Modeling
by: Liu, Shipeng, et al.
Published: (2026)
by: Liu, Shipeng, et al.
Published: (2026)
DirectSwap: Mask-Free Cross-Identity Training and Benchmarking for Expression-Consistent Video Head Swapping
by: Wang, Yanan, et al.
Published: (2025)
by: Wang, Yanan, et al.
Published: (2025)
VideoStudio: Generating Consistent-Content and Multi-Scene Videos
by: Long, Fuchen, et al.
Published: (2024)
by: Long, Fuchen, et al.
Published: (2024)
VideoSSR: Video Self-Supervised Reinforcement Learning
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
Similar Items
-
Semantic Segmentation on VSPW Dataset through Masked Video Consistency
by: Liang, Chen, et al.
Published: (2024) -
MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
by: Huang, Jun, et al.
Published: (2025) -
Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation
by: Lou, Zijie, et al.
Published: (2026) -
SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything Model
by: Yu, Chongkai, et al.
Published: (2024) -
GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing
by: Wang, Tong, et al.
Published: (2025)