Phased Consistency Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910726007816192 |
|---|---|
| author | Wang, Fu-Yun Huang, Zhaoyang Bergman, Alexander William Shen, Dazhong Gao, Peng Lingelbach, Michael Sun, Keqiang Bian, Weikang Song, Guanglu Liu, Yu Wang, Xiaogang Li, Hongsheng |
| author_facet | Wang, Fu-Yun Huang, Zhaoyang Bergman, Alexander William Shen, Dazhong Gao, Peng Lingelbach, Michael Sun, Keqiang Bian, Weikang Song, Guanglu Liu, Yu Wang, Xiaogang Li, Hongsheng |
| contents | Consistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of Latent Consistency Models (LCMs). We investigate the reasons behind these limitations and propose Phased Consistency Models (PCMs), which generalize the design space and address the identified limitations. Our evaluations demonstrate that PCMs outperform LCMs across 1--16 step generation settings. While PCMs are specifically designed for multi-step refinement, they achieve comparable 1-step generation results to previously state-of-the-art specifically designed 1-step methods. Furthermore, we show the methodology of PCMs is versatile and applicable to video generation, enabling us to train the state-of-the-art few-step text-to-video generator. Our code is available at https://github.com/G-U-N/Phased-Consistency-Model. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_18407 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Phased Consistency Models Wang, Fu-Yun Huang, Zhaoyang Bergman, Alexander William Shen, Dazhong Gao, Peng Lingelbach, Michael Sun, Keqiang Bian, Weikang Song, Guanglu Liu, Yu Wang, Xiaogang Li, Hongsheng Machine Learning Computer Vision and Pattern Recognition Consistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of Latent Consistency Models (LCMs). We investigate the reasons behind these limitations and propose Phased Consistency Models (PCMs), which generalize the design space and address the identified limitations. Our evaluations demonstrate that PCMs outperform LCMs across 1--16 step generation settings. While PCMs are specifically designed for multi-step refinement, they achieve comparable 1-step generation results to previously state-of-the-art specifically designed 1-step methods. Furthermore, we show the methodology of PCMs is versatile and applicable to video generation, enabling us to train the state-of-the-art few-step text-to-video generator. Our code is available at https://github.com/G-U-N/Phased-Consistency-Model. |
| title | Phased Consistency Models |
| topic | Machine Learning Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2405.18407 |