Laminating Representation Autoencoders for Efficient Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Calvo-González, Ramón, Fleuret, François
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917248312016896
author Calvo-González, Ramón
Fleuret, François
author_facet Calvo-González, Ramón
Fleuret, François
contents Recent work has shown that diffusion models can generate high-quality images by operating directly on SSL patch features rather than pixel-space latents. However, the dense patch grids from encoders like DINOv2 contain significant redundancy, making diffusion needlessly expensive. We introduce FlatDINO, a variational autoencoder that compresses this representation into a one-dimensional sequence of just 32 continuous tokens -an 8x reduction in sequence length and 48x compression in total dimensionality. On ImageNet 256x256, a DiT-XL trained on FlatDINO latents achieves a gFID of 1.80 with classifier-free guidance while requiring 8x fewer FLOPs per forward pass and up to 4.5x fewer FLOPs per training step compared to diffusion on uncompressed DINOv2 features. These are preliminary results and this work is in progress.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04873
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Laminating Representation Autoencoders for Efficient Diffusion
Calvo-González, Ramón
Fleuret, François
Computer Vision and Pattern Recognition
Recent work has shown that diffusion models can generate high-quality images by operating directly on SSL patch features rather than pixel-space latents. However, the dense patch grids from encoders like DINOv2 contain significant redundancy, making diffusion needlessly expensive. We introduce FlatDINO, a variational autoencoder that compresses this representation into a one-dimensional sequence of just 32 continuous tokens -an 8x reduction in sequence length and 48x compression in total dimensionality. On ImageNet 256x256, a DiT-XL trained on FlatDINO latents achieves a gFID of 1.80 with classifier-free guidance while requiring 8x fewer FLOPs per forward pass and up to 4.5x fewer FLOPs per training step compared to diffusion on uncompressed DINOv2 features. These are preliminary results and this work is in progress.
title Laminating Representation Autoencoders for Efficient Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.04873