VCoME: Verbal Video Composition with Multimodal Editing Effects

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gong, Weibo, Jin, Xiaojie, Li, Xin, He, Dongliang, Wu, Xinglong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914859662180352
author Gong, Weibo
Jin, Xiaojie
Li, Xin
He, Dongliang
Wu, Xinglong
author_facet Gong, Weibo
Jin, Xiaojie
Li, Xin
He, Dongliang
Wu, Xinglong
contents Verbal videos, featuring voice-overs or text overlays, provide valuable content but present significant challenges in composition, especially when incorporating editing effects to enhance clarity and visual appeal. In this paper, we introduce the novel task of verbal video composition with editing effects. This task aims to generate coherent and visually appealing verbal videos by integrating multimodal editing effects across textual, visual, and audio categories. To achieve this, we curate a large-scale dataset of video effects compositions from publicly available sources. We then formulate this task as a generative problem, involving the identification of appropriate positions in the verbal content and the recommendation of editing effects for these positions. To address this task, we propose VCoME, a general framework that employs a large multimodal model to generate editing effects for video composition. Specifically, VCoME takes in the multimodal video context and autoregressively outputs where to apply effects within the verbal content and which effects are most appropriate for each position. VCoME also supports prompt-based control of composition density and style, providing substantial flexibility for diverse applications. Through extensive quantitative and qualitative evaluations, we clearly demonstrate the effectiveness of VCoME. A comprehensive user study shows that our method produces videos of professional quality while being 85$\times$ more efficient than professional editors.
format Preprint
id arxiv_https___arxiv_org_abs_2407_04697
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VCoME: Verbal Video Composition with Multimodal Editing Effects
Gong, Weibo
Jin, Xiaojie
Li, Xin
He, Dongliang
Wu, Xinglong
Computer Vision and Pattern Recognition
Multimedia
Verbal videos, featuring voice-overs or text overlays, provide valuable content but present significant challenges in composition, especially when incorporating editing effects to enhance clarity and visual appeal. In this paper, we introduce the novel task of verbal video composition with editing effects. This task aims to generate coherent and visually appealing verbal videos by integrating multimodal editing effects across textual, visual, and audio categories. To achieve this, we curate a large-scale dataset of video effects compositions from publicly available sources. We then formulate this task as a generative problem, involving the identification of appropriate positions in the verbal content and the recommendation of editing effects for these positions. To address this task, we propose VCoME, a general framework that employs a large multimodal model to generate editing effects for video composition. Specifically, VCoME takes in the multimodal video context and autoregressively outputs where to apply effects within the verbal content and which effects are most appropriate for each position. VCoME also supports prompt-based control of composition density and style, providing substantial flexibility for diverse applications. Through extensive quantitative and qualitative evaluations, we clearly demonstrate the effectiveness of VCoME. A comprehensive user study shows that our method produces videos of professional quality while being 85$\times$ more efficient than professional editors.
title VCoME: Verbal Video Composition with Multimodal Editing Effects
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2407.04697