ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Tianming, Jiang, Haichao, Zheng, Wei-Shi, Hu, Jian-Fang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918037923299328
author Liang, Tianming
Jiang, Haichao
Zheng, Wei-Shi
Hu, Jian-Fang
author_facet Liang, Tianming
Jiang, Haichao
Zheng, Wei-Shi
Hu, Jian-Fang
contents Referring Video Object Segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This task has attracted increasing attention in the field of computer vision due to its promising applications in video editing and human-agent interaction. Recently, ReferDINO has demonstrated promising performance in this task by adapting object-level vision-language knowledge from pretrained foundational image models. In this report, we further enhance its capabilities by incorporating the advantages of SAM2 in mask quality and object consistency. In addition, to effectively balance performance between single-object and multi-object scenarios, we introduce a conditional mask fusion strategy that adaptively fuses the masks from ReferDINO and SAM2. Our solution, termed ReferDINO-Plus, achieves 60.43 \(\mathcal{J}\&\mathcal{F}\) on MeViS test set, securing 2nd place in the MeViS PVUW challenge at CVPR 2025. The code is available at: https://github.com/iSEE-Laboratory/ReferDINO-Plus.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23509
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025
Liang, Tianming
Jiang, Haichao
Zheng, Wei-Shi
Hu, Jian-Fang
Computer Vision and Pattern Recognition
Referring Video Object Segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This task has attracted increasing attention in the field of computer vision due to its promising applications in video editing and human-agent interaction. Recently, ReferDINO has demonstrated promising performance in this task by adapting object-level vision-language knowledge from pretrained foundational image models. In this report, we further enhance its capabilities by incorporating the advantages of SAM2 in mask quality and object consistency. In addition, to effectively balance performance between single-object and multi-object scenarios, we introduce a conditional mask fusion strategy that adaptively fuses the masks from ReferDINO and SAM2. Our solution, termed ReferDINO-Plus, achieves 60.43 \(\mathcal{J}\&\mathcal{F}\) on MeViS test set, securing 2nd place in the MeViS PVUW challenge at CVPR 2025. The code is available at: https://github.com/iSEE-Laboratory/ReferDINO-Plus.
title ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.23509