SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Nanye, Goldstein, Mark, Albergo, Michael S., Boffi, Nicholas M., Vanden-Eijnden, Eric, Xie, Saining
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917782458728448
author Ma, Nanye
Goldstein, Mark
Albergo, Michael S.
Boffi, Nicholas M.
Vanden-Eijnden, Eric
Xie, Saining
author_facet Ma, Nanye
Goldstein, Mark
Albergo, Michael S.
Boffi, Nicholas M.
Vanden-Eijnden, Eric
Xie, Saining
contents We present Scalable Interpolant Transformers (SiT), a family of generative models built on the backbone of Diffusion Transformers (DiT). The interpolant framework, which allows for connecting two distributions in a more flexible way than standard diffusion models, makes possible a modular study of various design choices impacting generative models built on dynamical transport: learning in discrete or continuous time, the objective function, the interpolant that connects the distributions, and deterministic or stochastic sampling. By carefully introducing the above ingredients, SiT surpasses DiT uniformly across model sizes on the conditional ImageNet 256x256 and 512x512 benchmark using the exact same model structure, number of parameters, and GFLOPs. By exploring various diffusion coefficients, which can be tuned separately from learning, SiT achieves an FID-50K score of 2.06 and 2.62, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08740
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
Ma, Nanye
Goldstein, Mark
Albergo, Michael S.
Boffi, Nicholas M.
Vanden-Eijnden, Eric
Xie, Saining
Computer Vision and Pattern Recognition
Machine Learning
We present Scalable Interpolant Transformers (SiT), a family of generative models built on the backbone of Diffusion Transformers (DiT). The interpolant framework, which allows for connecting two distributions in a more flexible way than standard diffusion models, makes possible a modular study of various design choices impacting generative models built on dynamical transport: learning in discrete or continuous time, the objective function, the interpolant that connects the distributions, and deterministic or stochastic sampling. By carefully introducing the above ingredients, SiT surpasses DiT uniformly across model sizes on the conditional ImageNet 256x256 and 512x512 benchmark using the exact same model structure, number of parameters, and GFLOPs. By exploring various diffusion coefficients, which can be tuned separately from learning, SiT achieves an FID-50K score of 2.06 and 2.62, respectively.
title SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2401.08740