MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Henghui, Liu, Chang, He, Shuting, Ying, Kaining, Jiang, Xudong, Loy, Chen Change, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
by: Ding, Henghui, et al.
Published: (2026)
by: Ding, Henghui, et al.
Published: (2026)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
by: Ying, Kaining, et al.
Published: (2025)
by: Ying, Kaining, et al.
Published: (2025)
MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes
by: Ding, Henghui, et al.
Published: (2025)
by: Ding, Henghui, et al.
Published: (2025)
Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation
by: He, Shuting, et al.
Published: (2024)
by: He, Shuting, et al.
Published: (2024)
MOVE: Motion-Guided Few-Shot Video Object Segmentation
by: Ying, Kaining, et al.
Published: (2025)
by: Ying, Kaining, et al.
Published: (2025)
Multimodal Referring Segmentation: A Survey
by: Ding, Henghui, et al.
Published: (2025)
by: Ding, Henghui, et al.
Published: (2025)
2nd Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation
by: Cao, Bin, et al.
Published: (2024)
by: Cao, Bin, et al.
Published: (2024)
Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
by: Wang, Yaxian, et al.
Published: (2025)
by: Wang, Yaxian, et al.
Published: (2025)
RefMask3D: Language-Guided Transformer for 3D Referring Segmentation
by: He, Shuting, et al.
Published: (2024)
by: He, Shuting, et al.
Published: (2024)
SegPoint: Segment Any Point Cloud via Large Language Model
by: He, Shuting, et al.
Published: (2024)
by: He, Shuting, et al.
Published: (2024)
3rd Place Solution for MeViS Track in CVPR 2024 PVUW workshop: Motion Expression guided Video Segmentation
by: Pan, Feiyu, et al.
Published: (2024)
by: Pan, Feiyu, et al.
Published: (2024)
1st Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
Segment Anything Across Shots: A Method and Benchmark
by: Hu, Hengrui, et al.
Published: (2025)
by: Hu, Hengrui, et al.
Published: (2025)
ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation
by: He, Xusheng, et al.
Published: (2026)
by: He, Xusheng, et al.
Published: (2026)
The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
by: Wang, Zhiyu, et al.
Published: (2026)
by: Wang, Zhiyu, et al.
Published: (2026)
Generalizable Implicit Motion Modeling for Video Frame Interpolation
by: Guo, Zujin, et al.
Published: (2024)
by: Guo, Zujin, et al.
Published: (2024)
ReferSplat: Referring Segmentation in 3D Gaussian Splatting
by: He, Shuting, et al.
Published: (2025)
by: He, Shuting, et al.
Published: (2025)
ROSE: Retrieval-Oriented Segmentation Enhancement
by: Tang, Song, et al.
Published: (2026)
by: Tang, Song, et al.
Published: (2026)
SAM3-DMS: Decoupled Memory Selection for Multi-target Video Segmentation of SAM3
by: Shen, Ruiqi, et al.
Published: (2026)
by: Shen, Ruiqi, et al.
Published: (2026)
4th PVUW MeViS 3rd Place Report: Sa2VA
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
3rd Place of MeViS-Audio Track of the 5th PVUW: VIRST-Audio
by: Hong, Jihwan, et al.
Published: (2026)
by: Hong, Jihwan, et al.
Published: (2026)
Explore In-Context Segmentation via Latent Diffusion Models
by: Wang, Chaoyang, et al.
Published: (2024)
by: Wang, Chaoyang, et al.
Published: (2024)
AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method
by: Miao, Deshui, et al.
Published: (2026)
by: Miao, Deshui, et al.
Published: (2026)
Evaluating SAM2 for Video Semantic Segmentation
by: Ariff, Syed Hesham Syed, et al.
Published: (2025)
by: Ariff, Syed Hesham Syed, et al.
Published: (2025)
Transformer-Based Visual Segmentation: A Survey
by: Li, Xiangtai, et al.
Published: (2023)
by: Li, Xiangtai, et al.
Published: (2023)
Mitigating the Curse of Dimensionality for Certified Robustness via Dual Randomized Smoothing
by: Xia, Song, et al.
Published: (2024)
by: Xia, Song, et al.
Published: (2024)
OMG-Seg: Is One Model Good Enough For All Segmentation?
by: Li, Xiangtai, et al.
Published: (2024)
by: Li, Xiangtai, et al.
Published: (2024)
Kalman-Inspired Feature Propagation for Video Face Super-Resolution
by: Feng, Ruicheng, et al.
Published: (2024)
by: Feng, Ruicheng, et al.
Published: (2024)
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
by: Sun, Ye, et al.
Published: (2025)
by: Sun, Ye, et al.
Published: (2025)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
Generate Your Talking Avatar from Video Reference
by: Guo, Zujin, et al.
Published: (2026)
by: Guo, Zujin, et al.
Published: (2026)
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
by: Gong, Dengxian, et al.
Published: (2026)
by: Gong, Dengxian, et al.
Published: (2026)
3D-GRES: Generalized 3D Referring Expression Segmentation
by: Wu, Changli, et al.
Published: (2024)
by: Wu, Changli, et al.
Published: (2024)
Open-set Anomaly Segmentation in Complex Scenarios
by: Xia, Song, et al.
Published: (2025)
by: Xia, Song, et al.
Published: (2025)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior
by: Chen, Honghua, et al.
Published: (2024)
by: Chen, Honghua, et al.
Published: (2024)
DifFace: Blind Face Restoration with Diffused Error Contraction
by: Yue, Zongsheng, et al.
Published: (2022)
by: Yue, Zongsheng, et al.
Published: (2022)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
by: He, Shuting, et al.
Published: (2025)
by: He, Shuting, et al.
Published: (2025)
Similar Items
-
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
by: Ding, Henghui, et al.
Published: (2026) -
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
by: Ying, Kaining, et al.
Published: (2025) -
MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes
by: Ding, Henghui, et al.
Published: (2025) -
Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation
by: He, Shuting, et al.
Published: (2024) -
MOVE: Motion-Guided Few-Shot Video Object Segmentation
by: Ying, Kaining, et al.
Published: (2025)