Spectrum Matching: a Unified Perspective for Superior Diffusability in Latent Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ning, Mang, Li, Mingxiao, Zhang, Le, Liu, Lanmiao, Blaschko, Matthew B., Salah, Albert Ali, Ertugrul, Itir Onal
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911518504779776
author Ning, Mang
Li, Mingxiao
Zhang, Le
Liu, Lanmiao
Blaschko, Matthew B.
Salah, Albert Ali
Ertugrul, Itir Onal
author_facet Ning, Mang
Li, Mingxiao
Zhang, Le
Liu, Lanmiao
Blaschko, Matthew B.
Salah, Albert Ali
Ertugrul, Itir Onal
contents In this paper, we study the diffusability (learnability) of variational autoencoders (VAE) in latent diffusion. First, we show that pixel-space diffusion trained with an MSE objective is inherently biased toward learning low and mid spatial frequencies, and that the power-law power spectral density (PSD) of natural images makes this bias perceptually beneficial. Motivated by this result, we propose the \emph{Spectrum Matching Hypothesis}: latents with superior diffusability should (i) follow a flattened power-law PSD (\emph{Encoding Spectrum Matching}, ESM) and (ii) preserve frequency-to-frequency semantic correspondence through the decoder (\emph{Decoding Spectrum Matching}, DSM). In practice, we apply ESM by matching the PSD between images and latents, and DSM via shared spectral masking with frequency-aligned reconstruction. Importantly, Spectrum Matching provides a unified view that clarifies prior observations of over-noisy or over-smoothed latents, and interprets several recent methods as special cases (e.g., VA-VAE, EQ-VAE). Experiments suggest that Spectrum Matching yields superior diffusion generation on CelebA and ImageNet datasets, and outperforms prior approaches. Finally, we extend the spectral view to representation alignment (REPA): we show that the directional spectral energy of the target representation is crucial for REPA, and propose a DoG-based method to further improve the performance of REPA. Our code is available https://github.com/forever208/SpectrumMatching.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14645
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Spectrum Matching: a Unified Perspective for Superior Diffusability in Latent Diffusion
Ning, Mang
Li, Mingxiao
Zhang, Le
Liu, Lanmiao
Blaschko, Matthew B.
Salah, Albert Ali
Ertugrul, Itir Onal
Computer Vision and Pattern Recognition
In this paper, we study the diffusability (learnability) of variational autoencoders (VAE) in latent diffusion. First, we show that pixel-space diffusion trained with an MSE objective is inherently biased toward learning low and mid spatial frequencies, and that the power-law power spectral density (PSD) of natural images makes this bias perceptually beneficial. Motivated by this result, we propose the \emph{Spectrum Matching Hypothesis}: latents with superior diffusability should (i) follow a flattened power-law PSD (\emph{Encoding Spectrum Matching}, ESM) and (ii) preserve frequency-to-frequency semantic correspondence through the decoder (\emph{Decoding Spectrum Matching}, DSM). In practice, we apply ESM by matching the PSD between images and latents, and DSM via shared spectral masking with frequency-aligned reconstruction. Importantly, Spectrum Matching provides a unified view that clarifies prior observations of over-noisy or over-smoothed latents, and interprets several recent methods as special cases (e.g., VA-VAE, EQ-VAE). Experiments suggest that Spectrum Matching yields superior diffusion generation on CelebA and ImageNet datasets, and outperforms prior approaches. Finally, we extend the spectral view to representation alignment (REPA): we show that the directional spectral energy of the target representation is crucial for REPA, and propose a DoG-based method to further improve the performance of REPA. Our code is available https://github.com/forever208/SpectrumMatching.
title Spectrum Matching: a Unified Perspective for Superior Diffusability in Latent Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.14645