AutoVFX: Physically Realistic Video Editing from Natural Language Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910683889664000 |
|---|---|
| author | Hsu, Hao-Yu Lin, Zhi-Hao Zhai, Albert Xia, Hongchi Wang, Shenlong |
| author_facet | Hsu, Hao-Yu Lin, Zhi-Hao Zhai, Albert Xia, Hongchi Wang, Shenlong |
| contents | Modern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and largely inaccessible to everyday users. In this work, we present AutoVFX, a framework that automatically creates realistic and dynamic VFX videos from a single video and natural language instructions. By carefully integrating neural scene modeling, LLM-based code generation, and physical simulation, AutoVFX is able to provide physically-grounded, photorealistic editing effects that can be controlled directly using natural language instructions. We conduct extensive experiments to validate AutoVFX's efficacy across a diverse spectrum of videos and instructions. Quantitative and qualitative results suggest that AutoVFX outperforms all competing methods by a large margin in generative quality, instruction alignment, editing versatility, and physical plausibility. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_02394 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AutoVFX: Physically Realistic Video Editing from Natural Language Instructions Hsu, Hao-Yu Lin, Zhi-Hao Zhai, Albert Xia, Hongchi Wang, Shenlong Computer Vision and Pattern Recognition Modern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and largely inaccessible to everyday users. In this work, we present AutoVFX, a framework that automatically creates realistic and dynamic VFX videos from a single video and natural language instructions. By carefully integrating neural scene modeling, LLM-based code generation, and physical simulation, AutoVFX is able to provide physically-grounded, photorealistic editing effects that can be controlled directly using natural language instructions. We conduct extensive experiments to validate AutoVFX's efficacy across a diverse spectrum of videos and instructions. Quantitative and qualitative results suggest that AutoVFX outperforms all competing methods by a large margin in generative quality, instruction alignment, editing versatility, and physical plausibility. |
| title | AutoVFX: Physically Realistic Video Editing from Natural Language Instructions |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2411.02394 |