GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lv, Jiaxi, Huang, Yi, Yan, Mingfu, Huang, Jiancheng, Liu, Jianzhuang, Liu, Yifan, Wen, Yafei, Chen, Xiaoxin, Chen, Shifeng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909178586464256
author Lv, Jiaxi
Huang, Yi
Yan, Mingfu
Huang, Jiancheng
Liu, Jianzhuang
Liu, Yifan
Wen, Yafei
Chen, Xiaoxin
Chen, Shifeng
author_facet Lv, Jiaxi
Huang, Yi
Yan, Mingfu
Huang, Jiancheng
Liu, Jianzhuang
Liu, Yifan
Wen, Yafei
Chen, Xiaoxin
Chen, Shifeng
contents Recent advances in text-to-video generation have harnessed the power of diffusion models to create visually compelling content conditioned on text prompts. However, they usually encounter high computational costs and often struggle to produce videos with coherent physical motions. To tackle these issues, we propose GPT4Motion, a training-free framework that leverages the planning capability of large language models such as GPT, the physical simulation strength of Blender, and the excellent image generation ability of text-to-image diffusion models to enhance the quality of video synthesis. Specifically, GPT4Motion employs GPT-4 to generate a Blender script based on a user textual prompt, which commands Blender's built-in physics engine to craft fundamental scene components that encapsulate coherent physical motions across frames. Then these components are inputted into Stable Diffusion to generate a video aligned with the textual prompt. Experimental results on three basic physical motion scenarios, including rigid object drop and collision, cloth draping and swinging, and liquid flow, demonstrate that GPT4Motion can generate high-quality videos efficiently in maintaining motion coherency and entity consistency. GPT4Motion offers new insights in text-to-video research, enhancing its quality and broadening its horizon for further explorations.
format Preprint
id arxiv_https___arxiv_org_abs_2311_12631
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning
Lv, Jiaxi
Huang, Yi
Yan, Mingfu
Huang, Jiancheng
Liu, Jianzhuang
Liu, Yifan
Wen, Yafei
Chen, Xiaoxin
Chen, Shifeng
Computer Vision and Pattern Recognition
Recent advances in text-to-video generation have harnessed the power of diffusion models to create visually compelling content conditioned on text prompts. However, they usually encounter high computational costs and often struggle to produce videos with coherent physical motions. To tackle these issues, we propose GPT4Motion, a training-free framework that leverages the planning capability of large language models such as GPT, the physical simulation strength of Blender, and the excellent image generation ability of text-to-image diffusion models to enhance the quality of video synthesis. Specifically, GPT4Motion employs GPT-4 to generate a Blender script based on a user textual prompt, which commands Blender's built-in physics engine to craft fundamental scene components that encapsulate coherent physical motions across frames. Then these components are inputted into Stable Diffusion to generate a video aligned with the textual prompt. Experimental results on three basic physical motion scenarios, including rigid object drop and collision, cloth draping and swinging, and liquid flow, demonstrate that GPT4Motion can generate high-quality videos efficiently in maintaining motion coherency and entity consistency. GPT4Motion offers new insights in text-to-video research, enhancing its quality and broadening its horizon for further explorations.
title GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.12631