Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yongliang, Zhu, Wenbo, Cao, Jiawang, Lu, Yi, Li, Bozheng, Chi, Weiheng, Qiu, Zihan, Su, Lirian, Zheng, Haolin, Wu, Jay, Yang, Xu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912157927473152
author Wu, Yongliang
Zhu, Wenbo
Cao, Jiawang
Lu, Yi
Li, Bozheng
Chi, Weiheng
Qiu, Zihan
Su, Lirian
Zheng, Haolin
Wu, Jay
Yang, Xu
author_facet Wu, Yongliang
Zhu, Wenbo
Cao, Jiawang
Lu, Yi
Li, Bozheng
Chi, Weiheng
Qiu, Zihan
Su, Lirian
Zheng, Haolin
Wu, Jay
Yang, Xu
contents The demand for producing short-form videos for sharing on social media platforms has experienced significant growth in recent times. Despite notable advancements in the fields of video summarization and highlight detection, which can create partially usable short films from raw videos, these approaches are often domain-specific and require an in-depth understanding of real-world video content. To tackle this predicament, we propose Repurpose-10K, an extensive dataset comprising over 10,000 videos with more than 120,000 annotated clips aimed at resolving the video long-to-short task. Recognizing the inherent constraints posed by untrained human annotators, which can result in inaccurate annotations for repurposed videos, we propose a two-stage solution to obtain annotations from real-world user-generated content. Furthermore, we offer a baseline model to address this challenging task by integrating audio, visual, and caption aspects through a cross-modal fusion and alignment framework. We aspire for our work to ignite groundbreaking research in the lesser-explored realms of video repurposing.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08879
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
Wu, Yongliang
Zhu, Wenbo
Cao, Jiawang
Lu, Yi
Li, Bozheng
Chi, Weiheng
Qiu, Zihan
Su, Lirian
Zheng, Haolin
Wu, Jay
Yang, Xu
Computer Vision and Pattern Recognition
The demand for producing short-form videos for sharing on social media platforms has experienced significant growth in recent times. Despite notable advancements in the fields of video summarization and highlight detection, which can create partially usable short films from raw videos, these approaches are often domain-specific and require an in-depth understanding of real-world video content. To tackle this predicament, we propose Repurpose-10K, an extensive dataset comprising over 10,000 videos with more than 120,000 annotated clips aimed at resolving the video long-to-short task. Recognizing the inherent constraints posed by untrained human annotators, which can result in inaccurate annotations for repurposed videos, we propose a two-stage solution to obtain annotations from real-world user-generated content. Furthermore, we offer a baseline model to address this challenging task by integrating audio, visual, and caption aspects through a cross-modal fusion and alignment framework. We aspire for our work to ignite groundbreaking research in the lesser-explored realms of video repurposing.
title Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.08879