Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918233625329664 |
|---|---|
| author | Liu, Shizhan Deng, Xinran Yang, Zhuoyi Teng, Jiayan Gu, Xiaotao Tang, Jie |
| author_facet | Liu, Shizhan Deng, Xinran Yang, Zhuoyi Teng, Jiayan Gu, Xiaotao Tang, Jie |
| contents | Latent diffusion models pair VAEs with diffusion backbones, and the structure of VAE latents strongly influences the difficulty of diffusion training. However, existing video VAEs typically focus on reconstruction fidelity, overlooking latent structure. We present a statistical analysis of video VAE latent spaces and identify two spectral properties essential for diffusion training: a spatio-temporal frequency spectrum biased toward low frequencies, and a channel-wise eigenspectrum dominated by a few modes. To induce these properties, we propose two lightweight, backbone-agnostic regularizers: Local Correlation Regularization and Latent Masked Reconstruction. Experiments show that our Spectral-Structured VAE (SSVAE) achieves a $3\times$ speedup in text-to-video generation convergence and a 10\% gain in video reward, outperforming strong open-source VAEs. The code is available at https://github.com/zai-org/SSVAE. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_05394 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability Liu, Shizhan Deng, Xinran Yang, Zhuoyi Teng, Jiayan Gu, Xiaotao Tang, Jie Computer Vision and Pattern Recognition Latent diffusion models pair VAEs with diffusion backbones, and the structure of VAE latents strongly influences the difficulty of diffusion training. However, existing video VAEs typically focus on reconstruction fidelity, overlooking latent structure. We present a statistical analysis of video VAE latent spaces and identify two spectral properties essential for diffusion training: a spatio-temporal frequency spectrum biased toward low frequencies, and a channel-wise eigenspectrum dominated by a few modes. To induce these properties, we propose two lightweight, backbone-agnostic regularizers: Local Correlation Regularization and Latent Masked Reconstruction. Experiments show that our Spectral-Structured VAE (SSVAE) achieves a $3\times$ speedup in text-to-video generation convergence and a 10\% gain in video reward, outperforming strong open-source VAEs. The code is available at https://github.com/zai-org/SSVAE. |
| title | Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.05394 |