Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Peize, Jiang, Yi, Chen, Shoufa, Zhang, Shilong, Peng, Bingyue, Luo, Ping, Yuan, Zehuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909220499095552
author Sun, Peize
Jiang, Yi
Chen, Shoufa
Zhang, Shilong
Peng, Bingyue
Luo, Ping
Yuan, Zehuan
author_facet Sun, Peize
Jiang, Yi
Chen, Shoufa
Zhang, Shilong
Peng, Bingyue
Luo, Ping
Yuan, Zehuan
contents We introduce LlamaGen, a new family of image generation models that apply original ``next-token prediction'' paradigm of large language models to visual generation domain. It is an affirmative answer to whether vanilla autoregressive models, e.g., Llama, without inductive biases on visual signals can achieve state-of-the-art image generation performance if scaling properly. We reexamine design spaces of image tokenizers, scalability properties of image generation models, and their training data quality. The outcome of this exploration consists of: (1) An image tokenizer with downsample ratio of 16, reconstruction quality of 0.94 rFID and codebook usage of 97% on ImageNet benchmark. (2) A series of class-conditional image generation models ranging from 111M to 3.1B parameters, achieving 2.18 FID on ImageNet 256x256 benchmarks, outperforming the popular diffusion models such as LDM, DiT. (3) A text-conditional image generation model with 775M parameters, from two-stage training on LAION-COCO and high aesthetics quality images, demonstrating competitive performance of visual quality and text alignment. (4) We verify the effectiveness of LLM serving frameworks in optimizing the inference speed of image generation models and achieve 326% - 414% speedup. We release all models and codes to facilitate open-source community of visual generation and multimodal foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06525
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Sun, Peize
Jiang, Yi
Chen, Shoufa
Zhang, Shilong
Peng, Bingyue
Luo, Ping
Yuan, Zehuan
Computer Vision and Pattern Recognition
We introduce LlamaGen, a new family of image generation models that apply original ``next-token prediction'' paradigm of large language models to visual generation domain. It is an affirmative answer to whether vanilla autoregressive models, e.g., Llama, without inductive biases on visual signals can achieve state-of-the-art image generation performance if scaling properly. We reexamine design spaces of image tokenizers, scalability properties of image generation models, and their training data quality. The outcome of this exploration consists of: (1) An image tokenizer with downsample ratio of 16, reconstruction quality of 0.94 rFID and codebook usage of 97% on ImageNet benchmark. (2) A series of class-conditional image generation models ranging from 111M to 3.1B parameters, achieving 2.18 FID on ImageNet 256x256 benchmarks, outperforming the popular diffusion models such as LDM, DiT. (3) A text-conditional image generation model with 775M parameters, from two-stage training on LAION-COCO and high aesthetics quality images, demonstrating competitive performance of visual quality and text alignment. (4) We verify the effectiveness of LLM serving frameworks in optimizing the inference speed of image generation models and achieve 326% - 414% speedup. We release all models and codes to facilitate open-source community of visual generation and multimodal foundation models.
title Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.06525