Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Shizhan, Deng, Xinran, Yang, Zhuoyi, Teng, Jiayan, Gu, Xiaotao, Tang, Jie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918233625329664
author Liu, Shizhan
Deng, Xinran
Yang, Zhuoyi
Teng, Jiayan
Gu, Xiaotao
Tang, Jie
author_facet Liu, Shizhan
Deng, Xinran
Yang, Zhuoyi
Teng, Jiayan
Gu, Xiaotao
Tang, Jie
contents Latent diffusion models pair VAEs with diffusion backbones, and the structure of VAE latents strongly influences the difficulty of diffusion training. However, existing video VAEs typically focus on reconstruction fidelity, overlooking latent structure. We present a statistical analysis of video VAE latent spaces and identify two spectral properties essential for diffusion training: a spatio-temporal frequency spectrum biased toward low frequencies, and a channel-wise eigenspectrum dominated by a few modes. To induce these properties, we propose two lightweight, backbone-agnostic regularizers: Local Correlation Regularization and Latent Masked Reconstruction. Experiments show that our Spectral-Structured VAE (SSVAE) achieves a $3\times$ speedup in text-to-video generation convergence and a 10\% gain in video reward, outperforming strong open-source VAEs. The code is available at https://github.com/zai-org/SSVAE.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05394
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
Liu, Shizhan
Deng, Xinran
Yang, Zhuoyi
Teng, Jiayan
Gu, Xiaotao
Tang, Jie
Computer Vision and Pattern Recognition
Latent diffusion models pair VAEs with diffusion backbones, and the structure of VAE latents strongly influences the difficulty of diffusion training. However, existing video VAEs typically focus on reconstruction fidelity, overlooking latent structure. We present a statistical analysis of video VAE latent spaces and identify two spectral properties essential for diffusion training: a spatio-temporal frequency spectrum biased toward low frequencies, and a channel-wise eigenspectrum dominated by a few modes. To induce these properties, we propose two lightweight, backbone-agnostic regularizers: Local Correlation Regularization and Latent Masked Reconstruction. Experiments show that our Spectral-Structured VAE (SSVAE) achieves a $3\times$ speedup in text-to-video generation convergence and a 10\% gain in video reward, outperforming strong open-source VAEs. The code is available at https://github.com/zai-org/SSVAE.
title Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05394