Autoregressive Image Generation without Vector Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Tianhong, Tian, Yonglong, Li, He, Deng, Mingyang, He, Kaiming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910679724720128
author Li, Tianhong
Tian, Yonglong
Li, He
Deng, Mingyang
He, Kaiming
author_facet Li, Tianhong
Tian, Yonglong
Li, He
Deng, Mingyang
He, Kaiming
contents Conventional wisdom holds that autoregressive models for image generation are typically accompanied by vector-quantized tokens. We observe that while a discrete-valued space can facilitate representing a categorical distribution, it is not a necessity for autoregressive modeling. In this work, we propose to model the per-token probability distribution using a diffusion procedure, which allows us to apply autoregressive models in a continuous-valued space. Rather than using categorical cross-entropy loss, we define a Diffusion Loss function to model the per-token probability. This approach eliminates the need for discrete-valued tokenizers. We evaluate its effectiveness across a wide range of cases, including standard autoregressive models and generalized masked autoregressive (MAR) variants. By removing vector quantization, our image generator achieves strong results while enjoying the speed advantage of sequence modeling. We hope this work will motivate the use of autoregressive generation in other continuous-valued domains and applications. Code is available at: https://github.com/LTH14/mar.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11838
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Autoregressive Image Generation without Vector Quantization
Li, Tianhong
Tian, Yonglong
Li, He
Deng, Mingyang
He, Kaiming
Computer Vision and Pattern Recognition
Conventional wisdom holds that autoregressive models for image generation are typically accompanied by vector-quantized tokens. We observe that while a discrete-valued space can facilitate representing a categorical distribution, it is not a necessity for autoregressive modeling. In this work, we propose to model the per-token probability distribution using a diffusion procedure, which allows us to apply autoregressive models in a continuous-valued space. Rather than using categorical cross-entropy loss, we define a Diffusion Loss function to model the per-token probability. This approach eliminates the need for discrete-valued tokenizers. We evaluate its effectiveness across a wide range of cases, including standard autoregressive models and generalized masked autoregressive (MAR) variants. By removing vector quantization, our image generator achieves strong results while enjoying the speed advantage of sequence modeling. We hope this work will motivate the use of autoregressive generation in other continuous-valued domains and applications. Code is available at: https://github.com/LTH14/mar.
title Autoregressive Image Generation without Vector Quantization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.11838