AutoVFX: Physically Realistic Video Editing from Natural Language Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsu, Hao-Yu, Lin, Zhi-Hao, Zhai, Albert, Xia, Hongchi, Wang, Shenlong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910683889664000
author Hsu, Hao-Yu
Lin, Zhi-Hao
Zhai, Albert
Xia, Hongchi
Wang, Shenlong
author_facet Hsu, Hao-Yu
Lin, Zhi-Hao
Zhai, Albert
Xia, Hongchi
Wang, Shenlong
contents Modern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and largely inaccessible to everyday users. In this work, we present AutoVFX, a framework that automatically creates realistic and dynamic VFX videos from a single video and natural language instructions. By carefully integrating neural scene modeling, LLM-based code generation, and physical simulation, AutoVFX is able to provide physically-grounded, photorealistic editing effects that can be controlled directly using natural language instructions. We conduct extensive experiments to validate AutoVFX's efficacy across a diverse spectrum of videos and instructions. Quantitative and qualitative results suggest that AutoVFX outperforms all competing methods by a large margin in generative quality, instruction alignment, editing versatility, and physical plausibility.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02394
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AutoVFX: Physically Realistic Video Editing from Natural Language Instructions
Hsu, Hao-Yu
Lin, Zhi-Hao
Zhai, Albert
Xia, Hongchi
Wang, Shenlong
Computer Vision and Pattern Recognition
Modern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and largely inaccessible to everyday users. In this work, we present AutoVFX, a framework that automatically creates realistic and dynamic VFX videos from a single video and natural language instructions. By carefully integrating neural scene modeling, LLM-based code generation, and physical simulation, AutoVFX is able to provide physically-grounded, photorealistic editing effects that can be controlled directly using natural language instructions. We conduct extensive experiments to validate AutoVFX's efficacy across a diverse spectrum of videos and instructions. Quantitative and qualitative results suggest that AutoVFX outperforms all competing methods by a large margin in generative quality, instruction alignment, editing versatility, and physical plausibility.
title AutoVFX: Physically Realistic Video Editing from Natural Language Instructions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.02394