Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhong, Yaokun, Jiang, Siyu, Zhu, Jian, Hu, Jian-Fang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916806618251264
author Zhong, Yaokun
Jiang, Siyu
Zhu, Jian
Hu, Jian-Fang
author_facet Zhong, Yaokun
Jiang, Siyu
Zhu, Jian
Hu, Jian-Fang
contents Semi-Supervised Video Paragraph Grounding (SSVPG) aims to localize multiple sentences in a paragraph from an untrimmed video with limited temporal annotations. Existing methods focus on teacher-student consistency learning and video-level contrastive loss, but they overlook the importance of perturbing query contexts to generate strong supervisory signals. In this work, we propose a novel Context Consistency Learning (CCL) framework that unifies the paradigms of consistency regularization and pseudo-labeling to enhance semi-supervised learning. Specifically, we first conduct teacher-student learning where the student model takes as inputs strongly-augmented samples with sentences removed and is enforced to learn from the adequately strong supervisory signals from the teacher model. Afterward, we conduct model retraining based on the generated pseudo labels, where the mutual agreement between the original and augmented views' predictions is utilized as the label confidence. Extensive experiments show that CCL outperforms existing methods by a large margin.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding
Zhong, Yaokun
Jiang, Siyu
Zhu, Jian
Hu, Jian-Fang
Computer Vision and Pattern Recognition
Semi-Supervised Video Paragraph Grounding (SSVPG) aims to localize multiple sentences in a paragraph from an untrimmed video with limited temporal annotations. Existing methods focus on teacher-student consistency learning and video-level contrastive loss, but they overlook the importance of perturbing query contexts to generate strong supervisory signals. In this work, we propose a novel Context Consistency Learning (CCL) framework that unifies the paradigms of consistency regularization and pseudo-labeling to enhance semi-supervised learning. Specifically, we first conduct teacher-student learning where the student model takes as inputs strongly-augmented samples with sentences removed and is enforced to learn from the adequately strong supervisory signals from the teacher model. Afterward, we conduct model retraining based on the generated pseudo labels, where the mutual agreement between the original and augmented views' predictions is utilized as the label confidence. Extensive experiments show that CCL outperforms existing methods by a large margin.
title Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.18476