Edit3K: Universal Representation Learning for Video Editing Components

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Xin, Zhang, Libo, Chen, Fan, Wen, Longyin, Wang, Yufei, Luo, Tiejian, Zhu, Sijie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917935397732352
author Gu, Xin
Zhang, Libo
Chen, Fan
Wen, Longyin
Wang, Yufei
Luo, Tiejian
Zhu, Sijie
author_facet Gu, Xin
Zhang, Libo
Chen, Fan
Wen, Longyin
Wang, Yufei
Luo, Tiejian
Zhu, Sijie
contents This paper focuses on understanding the predominant video creation pipeline, i.e., compositional video editing with six main types of editing components, including video effects, animation, transition, filter, sticker, and text. In contrast to existing visual representation learning of visual materials (i.e., images/videos), we aim to learn visual representations of editing actions/components that are generally applied on raw materials. We start by proposing the first large-scale dataset for editing components of video creation, which covers about $3,094$ editing components with $618,800$ videos. Each video in our dataset is rendered by various image/video materials with a single editing component, which supports atomic visual understanding of different editing components. It can also benefit several downstream tasks, e.g., editing component recommendation, editing component recognition/retrieval, etc. Existing visual representation methods perform poorly because it is difficult to disentangle the visual appearance of editing components from raw materials. To that end, we benchmark popular alternative solutions and propose a novel method that learns to attend to the appearance of editing components regardless of raw materials. Our method achieves favorable results on editing component retrieval/recognition compared to the alternative solutions. A user study is also conducted to show that our representations cluster visually similar editing components better than other alternatives. Furthermore, our learned representations used to transition recommendation tasks achieve state-of-the-art results on the AutoTransition dataset. The code and dataset are available at https://github.com/GX77/Edit3K .
format Preprint
id arxiv_https___arxiv_org_abs_2403_16048
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Edit3K: Universal Representation Learning for Video Editing Components
Gu, Xin
Zhang, Libo
Chen, Fan
Wen, Longyin
Wang, Yufei
Luo, Tiejian
Zhu, Sijie
Computer Vision and Pattern Recognition
This paper focuses on understanding the predominant video creation pipeline, i.e., compositional video editing with six main types of editing components, including video effects, animation, transition, filter, sticker, and text. In contrast to existing visual representation learning of visual materials (i.e., images/videos), we aim to learn visual representations of editing actions/components that are generally applied on raw materials. We start by proposing the first large-scale dataset for editing components of video creation, which covers about $3,094$ editing components with $618,800$ videos. Each video in our dataset is rendered by various image/video materials with a single editing component, which supports atomic visual understanding of different editing components. It can also benefit several downstream tasks, e.g., editing component recommendation, editing component recognition/retrieval, etc. Existing visual representation methods perform poorly because it is difficult to disentangle the visual appearance of editing components from raw materials. To that end, we benchmark popular alternative solutions and propose a novel method that learns to attend to the appearance of editing components regardless of raw materials. Our method achieves favorable results on editing component retrieval/recognition compared to the alternative solutions. A user study is also conducted to show that our representations cluster visually similar editing components better than other alternatives. Furthermore, our learned representations used to transition recommendation tasks achieve state-of-the-art results on the AutoTransition dataset. The code and dataset are available at https://github.com/GX77/Edit3K .
title Edit3K: Universal Representation Learning for Video Editing Components
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.16048