VideoXum: Cross-modal Visual and Textural Summarization of Videos

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Jingyang, Hua, Hang, Chen, Ming, Li, Yikang, Hsiao, Jenhao, Ho, Chiuman, Luo, Jiebo
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929324661145600
author Lin, Jingyang
Hua, Hang
Chen, Ming
Li, Yikang
Hsiao, Jenhao
Ho, Chiuman
Luo, Jiebo
author_facet Lin, Jingyang
Hua, Hang
Chen, Ming
Li, Yikang
Hsiao, Jenhao
Ho, Chiuman
Luo, Jiebo
contents Video summarization aims to distill the most important information from a source video to produce either an abridged clip or a textual narrative. Traditionally, different methods have been proposed depending on whether the output is a video or text, thus ignoring the correlation between the two semantically related tasks of visual summarization and textual summarization. We propose a new joint video and text summarization task. The goal is to generate both a shortened video clip along with the corresponding textual summary from a long video, collectively referred to as a cross-modal summary. The generated shortened video clip and text narratives should be semantically well aligned. To this end, we first build a large-scale human-annotated dataset -- VideoXum (X refers to different modalities). The dataset is reannotated based on ActivityNet. After we filter out the videos that do not meet the length requirements, 14,001 long videos remain in our new dataset. Each video in our reannotated dataset has human-annotated video summaries and the corresponding narrative summaries. We then design a novel end-to-end model -- VTSUM-BILP to address the challenges of our proposed task. Moreover, we propose a new metric called VT-CLIPScore to help evaluate the semantic consistency of cross-modality summary. The proposed model achieves promising performance on this new task and establishes a benchmark for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2303_12060
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VideoXum: Cross-modal Visual and Textural Summarization of Videos
Lin, Jingyang
Hua, Hang
Chen, Ming
Li, Yikang
Hsiao, Jenhao
Ho, Chiuman
Luo, Jiebo
Computer Vision and Pattern Recognition
Computation and Language
Video summarization aims to distill the most important information from a source video to produce either an abridged clip or a textual narrative. Traditionally, different methods have been proposed depending on whether the output is a video or text, thus ignoring the correlation between the two semantically related tasks of visual summarization and textual summarization. We propose a new joint video and text summarization task. The goal is to generate both a shortened video clip along with the corresponding textual summary from a long video, collectively referred to as a cross-modal summary. The generated shortened video clip and text narratives should be semantically well aligned. To this end, we first build a large-scale human-annotated dataset -- VideoXum (X refers to different modalities). The dataset is reannotated based on ActivityNet. After we filter out the videos that do not meet the length requirements, 14,001 long videos remain in our new dataset. Each video in our reannotated dataset has human-annotated video summaries and the corresponding narrative summaries. We then design a novel end-to-end model -- VTSUM-BILP to address the challenges of our proposed task. Moreover, we propose a new metric called VT-CLIPScore to help evaluate the semantic consistency of cross-modality summary. The proposed model achieves promising performance on this new task and establishes a benchmark for future research.
title VideoXum: Cross-modal Visual and Textural Summarization of Videos
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2303.12060