Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, David Junhao, Wu, Jay Zhangjie, Liu, Jia-Wei, Zhao, Rui, Ran, Lingmin, Gu, Yuchao, Gao, Difei, Shou, Mike Zheng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908384422264832
author Zhang, David Junhao
Wu, Jay Zhangjie
Liu, Jia-Wei
Zhao, Rui
Ran, Lingmin
Gu, Yuchao
Gao, Difei
Shou, Mike Zheng
author_facet Zhang, David Junhao
Wu, Jay Zhangjie
Liu, Jia-Wei
Zhao, Rui
Ran, Lingmin
Gu, Yuchao
Gao, Difei
Shou, Mike Zheng
contents Significant advancements have been achieved in the realm of large-scale pre-trained text-to-video Diffusion Models (VDMs). However, previous methods either rely solely on pixel-based VDMs, which come with high computational costs, or on latent-based VDMs, which often struggle with precise text-video alignment. In this paper, we are the first to propose a hybrid model, dubbed as Show-1, which marries pixel-based and latent-based VDMs for text-to-video generation. Our model first uses pixel-based VDMs to produce a low-resolution video of strong text-video correlation. After that, we propose a novel expert translation method that employs the latent-based VDMs to further upsample the low-resolution video to high resolution, which can also remove potential artifacts and corruptions from low-resolution videos. Compared to latent VDMs, Show-1 can produce high-quality videos of precise text-video alignment; Compared to pixel VDMs, Show-1 is much more efficient (GPU memory usage during inference is 15G vs 72G). Furthermore, our Show-1 model can be readily adapted for motion customization and video stylization applications through simple temporal attention layer finetuning. Our model achieves state-of-the-art performance on standard video generation benchmarks. Our code and model weights are publicly available at https://github.com/showlab/Show-1.
format Preprint
id arxiv_https___arxiv_org_abs_2309_15818
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
Zhang, David Junhao
Wu, Jay Zhangjie
Liu, Jia-Wei
Zhao, Rui
Ran, Lingmin
Gu, Yuchao
Gao, Difei
Shou, Mike Zheng
Computer Vision and Pattern Recognition
Significant advancements have been achieved in the realm of large-scale pre-trained text-to-video Diffusion Models (VDMs). However, previous methods either rely solely on pixel-based VDMs, which come with high computational costs, or on latent-based VDMs, which often struggle with precise text-video alignment. In this paper, we are the first to propose a hybrid model, dubbed as Show-1, which marries pixel-based and latent-based VDMs for text-to-video generation. Our model first uses pixel-based VDMs to produce a low-resolution video of strong text-video correlation. After that, we propose a novel expert translation method that employs the latent-based VDMs to further upsample the low-resolution video to high resolution, which can also remove potential artifacts and corruptions from low-resolution videos. Compared to latent VDMs, Show-1 can produce high-quality videos of precise text-video alignment; Compared to pixel VDMs, Show-1 is much more efficient (GPU memory usage during inference is 15G vs 72G). Furthermore, our Show-1 model can be readily adapted for motion customization and video stylization applications through simple temporal attention layer finetuning. Our model achieves state-of-the-art performance on standard video generation benchmarks. Our code and model weights are publicly available at https://github.com/showlab/Show-1.
title Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2309.15818