Video Summarization: Towards Entity-Aware Captions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ayyubi, Hammad A., Liu, Tianqi, Nagrani, Arsha, Lin, Xudong, Zhang, Mingda, Arnab, Anurag, Han, Feng, Zhu, Yukun, Liu, Jialu, Chang, Shih-Fu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912112450732032
author Ayyubi, Hammad A.
Liu, Tianqi
Nagrani, Arsha
Lin, Xudong
Zhang, Mingda
Arnab, Anurag
Han, Feng
Zhu, Yukun
Liu, Jialu
Chang, Shih-Fu
author_facet Ayyubi, Hammad A.
Liu, Tianqi
Nagrani, Arsha
Lin, Xudong
Zhang, Mingda
Arnab, Anurag
Han, Feng
Zhu, Yukun
Liu, Jialu
Chang, Shih-Fu
contents Existing popular video captioning benchmarks and models deal with generic captions devoid of specific person, place or organization named entities. In contrast, news videos present a challenging setting where the caption requires such named entities for meaningful summarization. As such, we propose the task of summarizing news video directly to entity-aware captions. We also release a large-scale dataset, VIEWS (VIdeo NEWS), to support research on this task. Further, we propose a method that augments visual information from videos with context retrieved from external world knowledge to generate entity-aware captions. We demonstrate the effectiveness of our approach on three video captioning models. We also show that our approach generalizes to existing news image captions dataset. With all the extensive experiments and insights, we believe we establish a solid basis for future research on this challenging task.
format Preprint
id arxiv_https___arxiv_org_abs_2312_02188
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Video Summarization: Towards Entity-Aware Captions
Ayyubi, Hammad A.
Liu, Tianqi
Nagrani, Arsha
Lin, Xudong
Zhang, Mingda
Arnab, Anurag
Han, Feng
Zhu, Yukun
Liu, Jialu
Chang, Shih-Fu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
Existing popular video captioning benchmarks and models deal with generic captions devoid of specific person, place or organization named entities. In contrast, news videos present a challenging setting where the caption requires such named entities for meaningful summarization. As such, we propose the task of summarizing news video directly to entity-aware captions. We also release a large-scale dataset, VIEWS (VIdeo NEWS), to support research on this task. Further, we propose a method that augments visual information from videos with context retrieved from external world knowledge to generate entity-aware captions. We demonstrate the effectiveness of our approach on three video captioning models. We also show that our approach generalizes to existing news image captions dataset. With all the extensive experiments and insights, we believe we establish a solid basis for future research on this challenging task.
title Video Summarization: Towards Entity-Aware Captions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
url https://arxiv.org/abs/2312.02188