Temporal Grounding as a Learning Signal for Referring Video Object Segmentation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lee, Seunghun, Seo, Jiwan, Kim, Jeonghoon, Moon, Sungho, Kim, Siwon, Yun, Haeun, Jeon, Hyogyeong, Choi, Wonhyeok, Jeong, Jaehoon, Durante, Zane, Park, Sang Hyun, Im, Sunghoon
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912614550863872
author Lee, Seunghun
Seo, Jiwan
Kim, Jeonghoon
Moon, Sungho
Kim, Siwon
Yun, Haeun
Jeon, Hyogyeong
Choi, Wonhyeok
Jeong, Jaehoon
Durante, Zane
Park, Sang Hyun
Im, Sunghoon
author_facet Lee, Seunghun
Seo, Jiwan
Kim, Jeonghoon
Moon, Sungho
Kim, Siwon
Yun, Haeun
Jeon, Hyogyeong
Choi, Wonhyeok
Jeong, Jaehoon
Durante, Zane
Park, Sang Hyun
Im, Sunghoon
contents Referring Video Object Segmentation (RVOS) aims to segment and track objects in videos based on natural language expressions, requiring precise alignment between visual content and textual queries. However, existing methods often suffer from semantic misalignment, largely due to indiscriminate frame sampling and supervision of all visible objects during training -- regardless of their actual relevance to the expression. We identify the core problem as the absence of an explicit temporal learning signal in conventional training paradigms. To address this, we introduce MeViS-M, a dataset built upon the challenging MeViS benchmark, where we manually annotate temporal spans when each object is referred to by the expression. These annotations provide a direct, semantically grounded supervision signal that was previously missing. To leverage this signal, we propose Temporally Grounded Learning (TGL), a novel learning framework that directly incorporates temporal grounding into the training process. Within this frame- work, we introduce two key strategies. First, Moment-guided Dual-path Propagation (MDP) improves both grounding and tracking by decoupling language-guided segmentation for relevant moments from language-agnostic propagation for others. Second, Object-level Selective Supervision (OSS) supervises only the objects temporally aligned with the expression in each training clip, thereby reducing semantic noise and reinforcing language-conditioned learning. Extensive experiments demonstrate that our TGL framework effectively leverages temporal signal to establish a new state-of-the-art on the challenging MeViS benchmark. We will make our code and the MeViS-M dataset publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11955
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
Lee, Seunghun
Seo, Jiwan
Kim, Jeonghoon
Moon, Sungho
Kim, Siwon
Yun, Haeun
Jeon, Hyogyeong
Choi, Wonhyeok
Jeong, Jaehoon
Durante, Zane
Park, Sang Hyun
Im, Sunghoon
Computer Vision and Pattern Recognition
Referring Video Object Segmentation (RVOS) aims to segment and track objects in videos based on natural language expressions, requiring precise alignment between visual content and textual queries. However, existing methods often suffer from semantic misalignment, largely due to indiscriminate frame sampling and supervision of all visible objects during training -- regardless of their actual relevance to the expression. We identify the core problem as the absence of an explicit temporal learning signal in conventional training paradigms. To address this, we introduce MeViS-M, a dataset built upon the challenging MeViS benchmark, where we manually annotate temporal spans when each object is referred to by the expression. These annotations provide a direct, semantically grounded supervision signal that was previously missing. To leverage this signal, we propose Temporally Grounded Learning (TGL), a novel learning framework that directly incorporates temporal grounding into the training process. Within this frame- work, we introduce two key strategies. First, Moment-guided Dual-path Propagation (MDP) improves both grounding and tracking by decoupling language-guided segmentation for relevant moments from language-agnostic propagation for others. Second, Object-level Selective Supervision (OSS) supervises only the objects temporally aligned with the expression in each training clip, thereby reducing semantic noise and reinforcing language-conditioned learning. Extensive experiments demonstrate that our TGL framework effectively leverages temporal signal to establish a new state-of-the-art on the challenging MeViS benchmark. We will make our code and the MeViS-M dataset publicly available.
title Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.11955