Global2Local: A Joint-Hierarchical Attention for Video Captioning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dai, Chengpeng, Chen, Fuhai, Sun, Xiaoshuai, Ji, Rongrong, Ye, Qixiang, Wu, Yongjian
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910909814800384
author Dai, Chengpeng
Chen, Fuhai
Sun, Xiaoshuai
Ji, Rongrong
Ye, Qixiang
Wu, Yongjian
author_facet Dai, Chengpeng
Chen, Fuhai
Sun, Xiaoshuai
Ji, Rongrong
Ye, Qixiang
Wu, Yongjian
contents Recently, automatic video captioning has attracted increasing attention, where the core challenge lies in capturing the key semantic items, like objects and actions as well as their spatial-temporal correlations from the redundant frames and semantic content. To this end, existing works select either the key video clips in a global level~(across multi frames), or key regions within each frame, which, however, neglect the hierarchical order, i.e., key frames first and key regions latter. In this paper, we propose a novel joint-hierarchical attention model for video captioning, which embeds the key clips, the key frames and the key regions jointly into the captioning model in a hierarchical manner. Such a joint-hierarchical attention model first conducts a global selection to identify key frames, followed by a Gumbel sampling operation to identify further key regions based on the key frames, achieving an accurate global-to-local feature representation to guide the captioning. Extensive quantitative evaluations on two public benchmark datasets MSVD and MSR-VTT demonstrates the superiority of the proposed method over the state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2203_06663
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Global2Local: A Joint-Hierarchical Attention for Video Captioning
Dai, Chengpeng
Chen, Fuhai
Sun, Xiaoshuai
Ji, Rongrong
Ye, Qixiang
Wu, Yongjian
Computer Vision and Pattern Recognition
Recently, automatic video captioning has attracted increasing attention, where the core challenge lies in capturing the key semantic items, like objects and actions as well as their spatial-temporal correlations from the redundant frames and semantic content. To this end, existing works select either the key video clips in a global level~(across multi frames), or key regions within each frame, which, however, neglect the hierarchical order, i.e., key frames first and key regions latter. In this paper, we propose a novel joint-hierarchical attention model for video captioning, which embeds the key clips, the key frames and the key regions jointly into the captioning model in a hierarchical manner. Such a joint-hierarchical attention model first conducts a global selection to identify key frames, followed by a Gumbel sampling operation to identify further key regions based on the key frames, achieving an accurate global-to-local feature representation to guide the captioning. Extensive quantitative evaluations on two public benchmark datasets MSVD and MSR-VTT demonstrates the superiority of the proposed method over the state-of-the-art methods.
title Global2Local: A Joint-Hierarchical Attention for Video Captioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2203.06663