MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Weimin, Liu, Jiawei, Lin, Zhijie, Yan, Jiangqiao, Chen, Shuo, Low, Chetwin, Hoang, Tuyen, Wu, Jie, Liew, Jun Hao, Yan, Hanshu, Zhou, Daquan, Feng, Jiashi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917562337460224
author Wang, Weimin
Liu, Jiawei
Lin, Zhijie
Yan, Jiangqiao
Chen, Shuo
Low, Chetwin
Hoang, Tuyen
Wu, Jie
Liew, Jun Hao
Yan, Hanshu
Zhou, Daquan
Feng, Jiashi
author_facet Wang, Weimin
Liu, Jiawei
Lin, Zhijie
Yan, Jiangqiao
Chen, Shuo
Low, Chetwin
Hoang, Tuyen
Wu, Jie
Liew, Jun Hao
Yan, Hanshu
Zhou, Daquan
Feng, Jiashi
contents The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that integrates the text-to-image model, video motion generator, reference image embedding module and frame interpolation module into an end-to-end video generation pipeline. Benefiting from these architecture designs, MagicVideo-V2 can generate an aesthetically pleasing, high-resolution video with remarkable fidelity and smoothness. It demonstrates superior performance over leading Text-to-Video systems such as Runway, Pika 1.0, Morph, Moon Valley and Stable Video Diffusion model via user evaluation at large scale.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04468
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
Wang, Weimin
Liu, Jiawei
Lin, Zhijie
Yan, Jiangqiao
Chen, Shuo
Low, Chetwin
Hoang, Tuyen
Wu, Jie
Liew, Jun Hao
Yan, Hanshu
Zhou, Daquan
Feng, Jiashi
Computer Vision and Pattern Recognition
Artificial Intelligence
The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that integrates the text-to-image model, video motion generator, reference image embedding module and frame interpolation module into an end-to-end video generation pipeline. Benefiting from these architecture designs, MagicVideo-V2 can generate an aesthetically pleasing, high-resolution video with remarkable fidelity and smoothness. It demonstrates superior performance over leading Text-to-Video systems such as Runway, Pika 1.0, Morph, Moon Valley and Stable Video Diffusion model via user evaluation at large scale.
title MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2401.04468