MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917562337460224 |
|---|---|
| author | Wang, Weimin Liu, Jiawei Lin, Zhijie Yan, Jiangqiao Chen, Shuo Low, Chetwin Hoang, Tuyen Wu, Jie Liew, Jun Hao Yan, Hanshu Zhou, Daquan Feng, Jiashi |
| author_facet | Wang, Weimin Liu, Jiawei Lin, Zhijie Yan, Jiangqiao Chen, Shuo Low, Chetwin Hoang, Tuyen Wu, Jie Liew, Jun Hao Yan, Hanshu Zhou, Daquan Feng, Jiashi |
| contents | The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that integrates the text-to-image model, video motion generator, reference image embedding module and frame interpolation module into an end-to-end video generation pipeline. Benefiting from these architecture designs, MagicVideo-V2 can generate an aesthetically pleasing, high-resolution video with remarkable fidelity and smoothness. It demonstrates superior performance over leading Text-to-Video systems such as Runway, Pika 1.0, Morph, Moon Valley and Stable Video Diffusion model via user evaluation at large scale. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_04468 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation Wang, Weimin Liu, Jiawei Lin, Zhijie Yan, Jiangqiao Chen, Shuo Low, Chetwin Hoang, Tuyen Wu, Jie Liew, Jun Hao Yan, Hanshu Zhou, Daquan Feng, Jiashi Computer Vision and Pattern Recognition Artificial Intelligence The growing demand for high-fidelity video generation from textual descriptions has catalyzed significant research in this field. In this work, we introduce MagicVideo-V2 that integrates the text-to-image model, video motion generator, reference image embedding module and frame interpolation module into an end-to-end video generation pipeline. Benefiting from these architecture designs, MagicVideo-V2 can generate an aesthetically pleasing, high-resolution video with remarkable fidelity and smoothness. It demonstrates superior performance over leading Text-to-Video systems such as Runway, Pika 1.0, Morph, Moon Valley and Stable Video Diffusion model via user evaluation at large scale. |
| title | MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2401.04468 |