SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Dongting, Chen, Jierun, Huang, Xijie, Coskun, Huseyin, Sahni, Arpit, Gupta, Aarush, Goyal, Anujraaj, Lahiri, Dishani, Singh, Rajesh, Idelbayev, Yerlan, Cao, Junli, Li, Yanyu, Cheng, Kwang-Ting, Chan, S. -H. Gary, Gong, Mingming, Tulyakov, Sergey, Kag, Anil, Xu, Yanwu, Ren, Jian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917867310546944
author Hu, Dongting
Chen, Jierun
Huang, Xijie
Coskun, Huseyin
Sahni, Arpit
Gupta, Aarush
Goyal, Anujraaj
Lahiri, Dishani
Singh, Rajesh
Idelbayev, Yerlan
Cao, Junli
Li, Yanyu
Cheng, Kwang-Ting
Chan, S. -H. Gary
Gong, Mingming
Tulyakov, Sergey
Kag, Anil
Xu, Yanwu
Ren, Jian
author_facet Hu, Dongting
Chen, Jierun
Huang, Xijie
Coskun, Huseyin
Sahni, Arpit
Gupta, Aarush
Goyal, Anujraaj
Lahiri, Dishani
Singh, Rajesh
Idelbayev, Yerlan
Cao, Junli
Li, Yanyu
Cheng, Kwang-Ting
Chan, S. -H. Gary
Gong, Mingming
Tulyakov, Sergey
Kag, Anil
Xu, Yanwu
Ren, Jian
contents Existing text-to-image (T2I) diffusion models face several limitations, including large model sizes, slow runtime, and low-quality generation on mobile devices. This paper aims to address all of these challenges by developing an extremely small and fast T2I model that generates high-resolution and high-quality images on mobile platforms. We propose several techniques to achieve this goal. First, we systematically examine the design choices of the network architecture to reduce model parameters and latency, while ensuring high-quality generation. Second, to further improve generation quality, we employ cross-architecture knowledge distillation from a much larger model, using a multi-level approach to guide the training of our model from scratch. Third, we enable a few-step generation by integrating adversarial guidance with knowledge distillation. For the first time, our model SnapGen, demonstrates the generation of 1024x1024 px images on a mobile device around 1.4 seconds. On ImageNet-1K, our model, with only 372M parameters, achieves an FID of 2.06 for 256x256 px generation. On T2I benchmarks (i.e., GenEval and DPG-Bench), our model with merely 379M parameters, surpasses large-scale models with billions of parameters at a significantly smaller size (e.g., 7x smaller than SDXL, 14x smaller than IF-XL).
format Preprint
id arxiv_https___arxiv_org_abs_2412_09619
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
Hu, Dongting
Chen, Jierun
Huang, Xijie
Coskun, Huseyin
Sahni, Arpit
Gupta, Aarush
Goyal, Anujraaj
Lahiri, Dishani
Singh, Rajesh
Idelbayev, Yerlan
Cao, Junli
Li, Yanyu
Cheng, Kwang-Ting
Chan, S. -H. Gary
Gong, Mingming
Tulyakov, Sergey
Kag, Anil
Xu, Yanwu
Ren, Jian
Computer Vision and Pattern Recognition
Existing text-to-image (T2I) diffusion models face several limitations, including large model sizes, slow runtime, and low-quality generation on mobile devices. This paper aims to address all of these challenges by developing an extremely small and fast T2I model that generates high-resolution and high-quality images on mobile platforms. We propose several techniques to achieve this goal. First, we systematically examine the design choices of the network architecture to reduce model parameters and latency, while ensuring high-quality generation. Second, to further improve generation quality, we employ cross-architecture knowledge distillation from a much larger model, using a multi-level approach to guide the training of our model from scratch. Third, we enable a few-step generation by integrating adversarial guidance with knowledge distillation. For the first time, our model SnapGen, demonstrates the generation of 1024x1024 px images on a mobile device around 1.4 seconds. On ImageNet-1K, our model, with only 372M parameters, achieves an FID of 2.06 for 256x256 px generation. On T2I benchmarks (i.e., GenEval and DPG-Bench), our model with merely 379M parameters, surpasses large-scale models with billions of parameters at a significantly smaller size (e.g., 7x smaller than SDXL, 14x smaller than IF-XL).
title SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09619