One Step Diffusion via Shortcut Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Frans, Kevin, Hafner, Danijar, Levine, Sergey, Abbeel, Pieter
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911019166597120
author Frans, Kevin
Hafner, Danijar
Levine, Sergey
Abbeel, Pieter
author_facet Frans, Kevin
Hafner, Danijar
Levine, Sergey
Abbeel, Pieter
contents Diffusion models and flow-matching models have enabled generating diverse and realistic images by learning to transfer noise to data. However, sampling from these models involves iterative denoising over many neural network passes, making generation slow and expensive. Previous approaches for speeding up sampling require complex training regimes, such as multiple training phases, multiple networks, or fragile scheduling. We introduce shortcut models, a family of generative models that use a single network and training phase to produce high-quality samples in a single or multiple sampling steps. Shortcut models condition the network not only on the current noise level but also on the desired step size, allowing the model to skip ahead in the generation process. Across a wide range of sampling step budgets, shortcut models consistently produce higher quality samples than previous approaches, such as consistency models and reflow. Compared to distillation, shortcut models reduce complexity to a single network and training phase and additionally allow varying step budgets at inference time.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12557
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle One Step Diffusion via Shortcut Models
Frans, Kevin
Hafner, Danijar
Levine, Sergey
Abbeel, Pieter
Machine Learning
Computer Vision and Pattern Recognition
Diffusion models and flow-matching models have enabled generating diverse and realistic images by learning to transfer noise to data. However, sampling from these models involves iterative denoising over many neural network passes, making generation slow and expensive. Previous approaches for speeding up sampling require complex training regimes, such as multiple training phases, multiple networks, or fragile scheduling. We introduce shortcut models, a family of generative models that use a single network and training phase to produce high-quality samples in a single or multiple sampling steps. Shortcut models condition the network not only on the current noise level but also on the desired step size, allowing the model to skip ahead in the generation process. Across a wide range of sampling step budgets, shortcut models consistently produce higher quality samples than previous approaches, such as consistency models and reflow. Compared to distillation, shortcut models reduce complexity to a single network and training phase and additionally allow varying step budgets at inference time.
title One Step Diffusion via Shortcut Models
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.12557