Causal Diffusion Transformers for Generative Modeling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Deng, Chaorui, Zhu, Deyao, Li, Kunchang, Guang, Shi, Fan, Haoqi
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909431377166336
author Deng, Chaorui
Zhu, Deyao
Li, Kunchang
Guang, Shi
Fan, Haoqi
author_facet Deng, Chaorui
Zhu, Deyao
Li, Kunchang
Guang, Shi
Fan, Haoqi
contents We introduce Causal Diffusion as the autoregressive (AR) counterpart of Diffusion models. It is a next-token(s) forecasting framework that is friendly to both discrete and continuous modalities and compatible with existing next-token prediction models like LLaMA and GPT. While recent works attempt to combine diffusion with AR models, we show that introducing sequential factorization to a diffusion model can substantially improve its performance and enables a smooth transition between AR and diffusion generation modes. Hence, we propose CausalFusion - a decoder-only transformer that dual-factorizes data across sequential tokens and diffusion noise levels, leading to state-of-the-art results on the ImageNet generation benchmark while also enjoying the AR advantage of generating an arbitrary number of tokens for in-context reasoning. We further demonstrate CausalFusion's multimodal capabilities through a joint image generation and captioning model, and showcase CausalFusion's ability for zero-shot in-context image manipulations. We hope that this work could provide the community with a fresh perspective on training multimodal models over discrete and continuous data.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12095
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Causal Diffusion Transformers for Generative Modeling
Deng, Chaorui
Zhu, Deyao
Li, Kunchang
Guang, Shi
Fan, Haoqi
Computer Vision and Pattern Recognition
We introduce Causal Diffusion as the autoregressive (AR) counterpart of Diffusion models. It is a next-token(s) forecasting framework that is friendly to both discrete and continuous modalities and compatible with existing next-token prediction models like LLaMA and GPT. While recent works attempt to combine diffusion with AR models, we show that introducing sequential factorization to a diffusion model can substantially improve its performance and enables a smooth transition between AR and diffusion generation modes. Hence, we propose CausalFusion - a decoder-only transformer that dual-factorizes data across sequential tokens and diffusion noise levels, leading to state-of-the-art results on the ImageNet generation benchmark while also enjoying the AR advantage of generating an arbitrary number of tokens for in-context reasoning. We further demonstrate CausalFusion's multimodal capabilities through a joint image generation and captioning model, and showcase CausalFusion's ability for zero-shot in-context image manipulations. We hope that this work could provide the community with a fresh perspective on training multimodal models over discrete and continuous data.
title Causal Diffusion Transformers for Generative Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.12095