Dense Video Captioning Using Unsupervised Semantic Information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Estevam, Valter, Laroca, Rayson, Pedrini, Helio, Menotti, David
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912176857415680
author Estevam, Valter
Laroca, Rayson
Pedrini, Helio
Menotti, David
author_facet Estevam, Valter
Laroca, Rayson
Pedrini, Helio
Menotti, David
contents We introduce a method to learn unsupervised semantic visual information based on the premise that complex events can be decomposed into simpler events and that these simple events are shared across several complex events. We first employ a clustering method to group representations producing a visual codebook. Then, we learn a dense representation by encoding the co-occurrence probability matrix for the codebook entries. This representation leverages the performance of the dense video captioning task in a scenario with only visual features. For example, we replace the audio signal in the BMT method and produce temporal proposals with comparable performance. Furthermore, we concatenate the visual representation with our descriptor in a vanilla transformer method to achieve state-of-the-art performance in the captioning subtask compared to the methods that explore only visual features, as well as a competitive performance with multi-modal methods. Our code is available at https://github.com/valterlej/dvcusi.
format Preprint
id arxiv_https___arxiv_org_abs_2112_08455
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Dense Video Captioning Using Unsupervised Semantic Information
Estevam, Valter
Laroca, Rayson
Pedrini, Helio
Menotti, David
Computer Vision and Pattern Recognition
We introduce a method to learn unsupervised semantic visual information based on the premise that complex events can be decomposed into simpler events and that these simple events are shared across several complex events. We first employ a clustering method to group representations producing a visual codebook. Then, we learn a dense representation by encoding the co-occurrence probability matrix for the codebook entries. This representation leverages the performance of the dense video captioning task in a scenario with only visual features. For example, we replace the audio signal in the BMT method and produce temporal proposals with comparable performance. Furthermore, we concatenate the visual representation with our descriptor in a vanilla transformer method to achieve state-of-the-art performance in the captioning subtask compared to the methods that explore only visual features, as well as a competitive performance with multi-modal methods. Our code is available at https://github.com/valterlej/dvcusi.
title Dense Video Captioning Using Unsupervised Semantic Information
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2112.08455