Normalizing Trajectory Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Gu, Jiatao, Chen, Tianrong, Shen, Ying, Berthelot, David, Zhai, Shuangfei, Susskind, Josh
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910213584453632
author Gu, Jiatao
Chen, Tianrong
Shen, Ying
Berthelot, David
Zhai, Shuangfei
Susskind, Josh
author_facet Gu, Jiatao
Chen, Tianrong
Shen, Ying
Berthelot, David
Zhai, Shuangfei
Susskind, Josh
contents Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a few coarse transitions. Existing few-step methods address this through distillation, consistency training, or adversarial objectives, but sacrifice the likelihood framework in the process. We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training. Architecturally, NTM combines shallow invertible blocks within each step with a deep parallel predictor across the trajectory, forming an end-to-end network trainable from scratch or initializable from pretrained flow-matching models. Its exact trajectory likelihood further enables self-distillation: a lightweight denoiser trained on the model's own score produces high-quality samples in four steps. On text-to-image benchmarks, NTM matches or outperforms strong image generation baselines in just four sampling steps while uniquely retaining exact likelihood over the generative trajectory.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08078
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Normalizing Trajectory Models
Gu, Jiatao
Chen, Tianrong
Shen, Ying
Berthelot, David
Zhai, Shuangfei
Susskind, Josh
Computer Vision and Pattern Recognition
Machine Learning
Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a few coarse transitions. Existing few-step methods address this through distillation, consistency training, or adversarial objectives, but sacrifice the likelihood framework in the process. We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training. Architecturally, NTM combines shallow invertible blocks within each step with a deep parallel predictor across the trajectory, forming an end-to-end network trainable from scratch or initializable from pretrained flow-matching models. Its exact trajectory likelihood further enables self-distillation: a lightweight denoiser trained on the model's own score produces high-quality samples in four steps. On text-to-image benchmarks, NTM matches or outperforms strong image generation baselines in just four sampling steps while uniquely retaining exact likelihood over the generative trajectory.
title Normalizing Trajectory Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.08078