Phased Consistency Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Fu-Yun, Huang, Zhaoyang, Bergman, Alexander William, Shen, Dazhong, Gao, Peng, Lingelbach, Michael, Sun, Keqiang, Bian, Weikang, Song, Guanglu, Liu, Yu, Wang, Xiaogang, Li, Hongsheng
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910726007816192
author Wang, Fu-Yun
Huang, Zhaoyang
Bergman, Alexander William
Shen, Dazhong
Gao, Peng
Lingelbach, Michael
Sun, Keqiang
Bian, Weikang
Song, Guanglu
Liu, Yu
Wang, Xiaogang
Li, Hongsheng
author_facet Wang, Fu-Yun
Huang, Zhaoyang
Bergman, Alexander William
Shen, Dazhong
Gao, Peng
Lingelbach, Michael
Sun, Keqiang
Bian, Weikang
Song, Guanglu
Liu, Yu
Wang, Xiaogang
Li, Hongsheng
contents Consistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of Latent Consistency Models (LCMs). We investigate the reasons behind these limitations and propose Phased Consistency Models (PCMs), which generalize the design space and address the identified limitations. Our evaluations demonstrate that PCMs outperform LCMs across 1--16 step generation settings. While PCMs are specifically designed for multi-step refinement, they achieve comparable 1-step generation results to previously state-of-the-art specifically designed 1-step methods. Furthermore, we show the methodology of PCMs is versatile and applicable to video generation, enabling us to train the state-of-the-art few-step text-to-video generator. Our code is available at https://github.com/G-U-N/Phased-Consistency-Model.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18407
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Phased Consistency Models
Wang, Fu-Yun
Huang, Zhaoyang
Bergman, Alexander William
Shen, Dazhong
Gao, Peng
Lingelbach, Michael
Sun, Keqiang
Bian, Weikang
Song, Guanglu
Liu, Yu
Wang, Xiaogang
Li, Hongsheng
Machine Learning
Computer Vision and Pattern Recognition
Consistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of Latent Consistency Models (LCMs). We investigate the reasons behind these limitations and propose Phased Consistency Models (PCMs), which generalize the design space and address the identified limitations. Our evaluations demonstrate that PCMs outperform LCMs across 1--16 step generation settings. While PCMs are specifically designed for multi-step refinement, they achieve comparable 1-step generation results to previously state-of-the-art specifically designed 1-step methods. Furthermore, we show the methodology of PCMs is versatile and applicable to video generation, enabling us to train the state-of-the-art few-step text-to-video generator. Our code is available at https://github.com/G-U-N/Phased-Consistency-Model.
title Phased Consistency Models
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.18407