Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Sun-Hyuk, Jo, Hayoung, Lee, Seong-Whan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909456918380544
author Choi, Sun-Hyuk
Jo, Hayoung
Lee, Seong-Whan
author_facet Choi, Sun-Hyuk
Jo, Hayoung
Lee, Seong-Whan
contents Referring video object segmentation aims to segment objects within a video corresponding to a given text description. Existing transformer-based temporal modeling approaches face challenges related to query inconsistency and the limited consideration of context. Query inconsistency produces unstable masks of different objects in the middle of the video. The limited consideration of context leads to the segmentation of incorrect objects by failing to adequately account for the relationship between the given text and instances. To address these issues, we propose the Multi-context Temporal Consistency Module (MTCM), which consists of an Aligner and a Multi-Context Enhancer (MCE). The Aligner removes noise from queries and aligns them to achieve query consistency. The MCE predicts text-relevant queries by considering multi-context. We applied MTCM to four different models, increasing performance across all of them, particularly achieving 47.6 J&F on the MeViS. Code is available at https://github.com/Choi58/MTCM.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04939
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation
Choi, Sun-Hyuk
Jo, Hayoung
Lee, Seong-Whan
Computer Vision and Pattern Recognition
Referring video object segmentation aims to segment objects within a video corresponding to a given text description. Existing transformer-based temporal modeling approaches face challenges related to query inconsistency and the limited consideration of context. Query inconsistency produces unstable masks of different objects in the middle of the video. The limited consideration of context leads to the segmentation of incorrect objects by failing to adequately account for the relationship between the given text and instances. To address these issues, we propose the Multi-context Temporal Consistency Module (MTCM), which consists of an Aligner and a Multi-Context Enhancer (MCE). The Aligner removes noise from queries and aligns them to achieve query consistency. The MCE predicts text-relevant queries by considering multi-context. We applied MTCM to four different models, increasing performance across all of them, particularly achieving 47.6 J&F on the MeViS. Code is available at https://github.com/Choi58/MTCM.
title Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04939