Scaling Up Video Summarization Pretraining with Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Argaw, Dawit Mureja, Yoon, Seunghyun, Heilbron, Fabian Caba, Deilamsalehy, Hanieh, Bui, Trung, Wang, Zhaowen, Dernoncourt, Franck, Chung, Joon Son
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929303285923840
author Argaw, Dawit Mureja
Yoon, Seunghyun
Heilbron, Fabian Caba
Deilamsalehy, Hanieh
Bui, Trung
Wang, Zhaowen
Dernoncourt, Franck
Chung, Joon Son
author_facet Argaw, Dawit Mureja
Yoon, Seunghyun
Heilbron, Fabian Caba
Deilamsalehy, Hanieh
Bui, Trung
Wang, Zhaowen
Dernoncourt, Franck
Chung, Joon Son
contents Long-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However, existing video summarization datasets are notably limited in their size, constraining the effectiveness of state-of-the-art methods for generalization. Our work aims to overcome this limitation by capitalizing on the abundance of long-form videos with dense speech-to-video alignment and the remarkable capabilities of recent large language models (LLMs) in summarizing long text. We introduce an automated and scalable pipeline for generating a large-scale video summarization dataset using LLMs as Oracle summarizers. By leveraging the generated dataset, we analyze the limitations of existing approaches and propose a new video summarization model that effectively addresses them. To facilitate further research in the field, our work also presents a new benchmark dataset that contains 1200 long videos each with high-quality summaries annotated by professionals. Extensive experiments clearly indicate that our proposed approach sets a new state-of-the-art in video summarization across several benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03398
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Up Video Summarization Pretraining with Large Language Models
Argaw, Dawit Mureja
Yoon, Seunghyun
Heilbron, Fabian Caba
Deilamsalehy, Hanieh
Bui, Trung
Wang, Zhaowen
Dernoncourt, Franck
Chung, Joon Son
Computer Vision and Pattern Recognition
Long-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However, existing video summarization datasets are notably limited in their size, constraining the effectiveness of state-of-the-art methods for generalization. Our work aims to overcome this limitation by capitalizing on the abundance of long-form videos with dense speech-to-video alignment and the remarkable capabilities of recent large language models (LLMs) in summarizing long text. We introduce an automated and scalable pipeline for generating a large-scale video summarization dataset using LLMs as Oracle summarizers. By leveraging the generated dataset, we analyze the limitations of existing approaches and propose a new video summarization model that effectively addresses them. To facilitate further research in the field, our work also presents a new benchmark dataset that contains 1200 long videos each with high-quality summaries annotated by professionals. Extensive experiments clearly indicate that our proposed approach sets a new state-of-the-art in video summarization across several benchmarks.
title Scaling Up Video Summarization Pretraining with Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.03398