Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Zhuyang, Yang, Yan, Yu, Yankai, Wang, Jie, Jiang, Yongquan, Wu, Xiao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913613704331264
author Xie, Zhuyang
Yang, Yan
Yu, Yankai
Wang, Jie
Jiang, Yongquan
Wu, Xiao
author_facet Xie, Zhuyang
Yang, Yan
Yu, Yankai
Wang, Jie
Jiang, Yongquan
Wu, Xiao
contents Dense video captioning aims to detect and describe all events in untrimmed videos. This paper presents a dense video captioning network called Multi-Concept Cyclic Learning (MCCL), which aims to: (1) detect multiple concepts at the frame level, using these concepts to enhance video features and provide temporal event cues; and (2) design cyclic co-learning between the generator and the localizer within the captioning network to promote semantic perception and event localization. Specifically, we perform weakly supervised concept detection for each frame, and the detected concept embeddings are integrated into the video features to provide event cues. Additionally, video-level concept contrastive learning is introduced to obtain more discriminative concept embeddings. In the captioning network, we establish a cyclic co-learning strategy where the generator guides the localizer for event localization through semantic matching, while the localizer enhances the generator's event semantic perception through location matching, making semantic perception and event localization mutually beneficial. MCCL achieves state-of-the-art performance on the ActivityNet Captions and YouCook2 datasets. Extensive experiments demonstrate its effectiveness and interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11467
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning
Xie, Zhuyang
Yang, Yan
Yu, Yankai
Wang, Jie
Jiang, Yongquan
Wu, Xiao
Computer Vision and Pattern Recognition
Dense video captioning aims to detect and describe all events in untrimmed videos. This paper presents a dense video captioning network called Multi-Concept Cyclic Learning (MCCL), which aims to: (1) detect multiple concepts at the frame level, using these concepts to enhance video features and provide temporal event cues; and (2) design cyclic co-learning between the generator and the localizer within the captioning network to promote semantic perception and event localization. Specifically, we perform weakly supervised concept detection for each frame, and the detected concept embeddings are integrated into the video features to provide event cues. Additionally, video-level concept contrastive learning is introduced to obtain more discriminative concept embeddings. In the captioning network, we establish a cyclic co-learning strategy where the generator guides the localizer for event localization through semantic matching, while the localizer enhances the generator's event semantic perception through location matching, making semantic perception and event localization mutually beneficial. MCCL achieves state-of-the-art performance on the ActivityNet Captions and YouCook2 datasets. Extensive experiments demonstrate its effectiveness and interpretability.
title Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.11467