Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kilian, Maciej, Jampani, Varun, Zettlemoyer, Luke
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910458565361664
author Kilian, Maciej
Jampani, Varun
Zettlemoyer, Luke
author_facet Kilian, Maciej
Jampani, Varun
Zettlemoyer, Luke
contents Nearly every recent image synthesis approach, including diffusion, masked-token prediction, and next-token prediction, uses a Transformer network architecture. Despite this common backbone, there has been no direct, compute controlled comparison of how these approaches affect performance and efficiency. We analyze the scalability of each approach through the lens of compute budget measured in FLOPs. We find that token prediction methods, led by next-token prediction, significantly outperform diffusion on prompt following. On image quality, while next-token prediction initially performs better, scaling trends suggest it is eventually matched by diffusion. We compare the inference compute efficiency of each approach and find that next token prediction is by far the most efficient. Based on our findings we recommend diffusion for applications targeting image quality and low latency; and next-token prediction when prompt following or throughput is more important.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13218
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
Kilian, Maciej
Jampani, Varun
Zettlemoyer, Luke
Computer Vision and Pattern Recognition
Nearly every recent image synthesis approach, including diffusion, masked-token prediction, and next-token prediction, uses a Transformer network architecture. Despite this common backbone, there has been no direct, compute controlled comparison of how these approaches affect performance and efficiency. We analyze the scalability of each approach through the lens of compute budget measured in FLOPs. We find that token prediction methods, led by next-token prediction, significantly outperform diffusion on prompt following. On image quality, while next-token prediction initially performs better, scaling trends suggest it is eventually matched by diffusion. We compare the inference compute efficiency of each approach and find that next token prediction is by far the most efficient. Based on our findings we recommend diffusion for applications targeting image quality and low latency; and next-token prediction when prompt following or throughput is more important.
title Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.13218