REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Weihan, Ma, Yimeng, Huang, Jingyue, Li, Yang, Ma, Wenye, Berg-Kirkpatrick, Taylor, McAuley, Julian, Liang, Paul Pu, Dong, Hao-Wen
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910967307173888
author Xu, Weihan
Ma, Yimeng
Huang, Jingyue
Li, Yang
Ma, Wenye
Berg-Kirkpatrick, Taylor
McAuley, Julian
Liang, Paul Pu
Dong, Hao-Wen
author_facet Xu, Weihan
Ma, Yimeng
Huang, Jingyue
Li, Yang
Ma, Wenye
Berg-Kirkpatrick, Taylor
McAuley, Julian
Liang, Paul Pu
Dong, Hao-Wen
contents Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive methods cannot `quote' from the input videos, i.e., inserting short video clips in their outputs. In this work, we explore novel video editing models for generating shorts that feature a coherent narrative with embedded video insertions extracted from a long input video. We propose a novel retrieval-embedded generation framework that allows a large language model to quote multimodal resources while maintaining a coherent narrative. Our proposed REGen system first generates the output story script with quote placeholders using a finetuned large language model, and then uses a novel retrieval model to replace the quote placeholders by selecting a video clip that best supports the narrative from a pool of candidate quotable video clips. We examine the proposed method on the task of documentary teaser generation, where short interview insertions are commonly used to support the narrative of a documentary. Our objective evaluations show that the proposed method can effectively insert short video clips while maintaining a coherent narrative. In a subjective survey, we show that our proposed method outperforms existing abstractive and extractive approaches in terms of coherence, alignment, and realism in teaser generation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18880
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
Xu, Weihan
Ma, Yimeng
Huang, Jingyue
Li, Yang
Ma, Wenye
Berg-Kirkpatrick, Taylor
McAuley, Julian
Liang, Paul Pu
Dong, Hao-Wen
Computer Vision and Pattern Recognition
Artificial Intelligence
Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive methods cannot `quote' from the input videos, i.e., inserting short video clips in their outputs. In this work, we explore novel video editing models for generating shorts that feature a coherent narrative with embedded video insertions extracted from a long input video. We propose a novel retrieval-embedded generation framework that allows a large language model to quote multimodal resources while maintaining a coherent narrative. Our proposed REGen system first generates the output story script with quote placeholders using a finetuned large language model, and then uses a novel retrieval model to replace the quote placeholders by selecting a video clip that best supports the narrative from a pool of candidate quotable video clips. We examine the proposed method on the task of documentary teaser generation, where short interview insertions are commonly used to support the narrative of a documentary. Our objective evaluations show that the proposed method can effectively insert short video clips while maintaining a coherent narrative. In a subjective survey, we show that our proposed method outperforms existing abstractive and extractive approaches in terms of coherence, alignment, and realism in teaser generation.
title REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.18880