Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shizhan, Deng, Xinran, Yang, Zhuoyi, Teng, Jiayan, Gu, Xiaotao, Tang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Concat-ID: Towards Universal Identity-Preserving Video Synthesis
by: Zhong, Yong, et al.
Published: (2025)
by: Zhong, Yong, et al.
Published: (2025)
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
by: Zheng, Wendi, et al.
Published: (2024)
by: Zheng, Wendi, et al.
Published: (2024)
Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
by: Zhang, Zhenxing, et al.
Published: (2025)
by: Zhang, Zhenxing, et al.
Published: (2025)
SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
by: Yan, Wenhao, et al.
Published: (2025)
by: Yan, Wenhao, et al.
Published: (2025)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
Delving into Spectral Clustering with Vision-Language Representations
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025)
by: Hong, Wenyi, et al.
Published: (2025)
RGB Pre-Training Enhanced Unobservable Feature Latent Diffusion Model for Spectral Reconstruction
by: Deng, Keli, et al.
Published: (2025)
by: Deng, Keli, et al.
Published: (2025)
VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
by: Xu, Jiazheng, et al.
Published: (2024)
by: Xu, Jiazheng, et al.
Published: (2024)
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
by: Xu, Sicheng, et al.
Published: (2026)
by: Xu, Sicheng, et al.
Published: (2026)
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
by: Zou, Ya, et al.
Published: (2025)
by: Zou, Ya, et al.
Published: (2025)
Seer: Language Instructed Video Prediction with Latent Diffusion Models
by: Gu, Xianfan, et al.
Published: (2023)
by: Gu, Xianfan, et al.
Published: (2023)
ACCORD: Alleviating Concept Coupling through Dependence Regularization for Text-to-Image Diffusion Personalization
by: Liu, Shizhan, et al.
Published: (2025)
by: Liu, Shizhan, et al.
Published: (2025)
LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
by: Chen, Qi, et al.
Published: (2026)
by: Chen, Qi, et al.
Published: (2026)
Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video
by: Shi, Yahao, et al.
Published: (2025)
by: Shi, Yahao, et al.
Published: (2025)
Spectrum Matching: a Unified Perspective for Superior Diffusability in Latent Diffusion
by: Ning, Mang, et al.
Published: (2026)
by: Ning, Mang, et al.
Published: (2026)
Delving into Mapping Uncertainty for Mapless Trajectory Prediction
by: Zhang, Zongzheng, et al.
Published: (2025)
by: Zhang, Zongzheng, et al.
Published: (2025)
CLIPSym: Delving into Symmetry Detection with CLIP
by: Yang, Tinghan, et al.
Published: (2025)
by: Yang, Tinghan, et al.
Published: (2025)
Delving Deep into Semantic Relation Distillation
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
CODA: Repurposing Continuous VAEs for Discrete Tokenization
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
Training-Free Vector Quantization via Gaussian VAEs
by: Xu, Tongda, et al.
Published: (2025)
by: Xu, Tongda, et al.
Published: (2025)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
by: Zhou, Zhenghong, et al.
Published: (2024)
by: Zhou, Zhenghong, et al.
Published: (2024)
Learning Energy-based Variational Latent Prior for VAEs
by: Dutta, Debottam, et al.
Published: (2025)
by: Dutta, Debottam, et al.
Published: (2025)
LTX-Video: Realtime Video Latent Diffusion
by: HaCohen, Yoav, et al.
Published: (2024)
by: HaCohen, Yoav, et al.
Published: (2024)
Adaptive 1D Video Diffusion Autoencoder
by: Teng, Yao, et al.
Published: (2026)
by: Teng, Yao, et al.
Published: (2026)
Latte: Latent Diffusion Transformer for Video Generation
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
On Inductive Biases That Enable Generalization of Diffusion Transformers
by: An, Jie, et al.
Published: (2024)
by: An, Jie, et al.
Published: (2024)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
by: Zhang, David Junhao, et al.
Published: (2023)
by: Zhang, David Junhao, et al.
Published: (2023)
PR-MIM: Delving Deeper into Partial Reconstruction in Masked Image Modeling
by: Li, Zhong-Yu, et al.
Published: (2024)
by: Li, Zhong-Yu, et al.
Published: (2024)
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
by: Gu, Bo, et al.
Published: (2026)
by: Gu, Bo, et al.
Published: (2026)
LatentColorization: Latent Diffusion-Based Speaker Video Colorization
by: Ward, Rory, et al.
Published: (2024)
by: Ward, Rory, et al.
Published: (2024)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
LVMark: Robust Watermark for Latent Video Diffusion Models
by: Jang, MinHyuk, et al.
Published: (2024)
by: Jang, MinHyuk, et al.
Published: (2024)
Motion-aware Latent Diffusion Models for Video Frame Interpolation
by: Huang, Zhilin, et al.
Published: (2024)
by: Huang, Zhilin, et al.
Published: (2024)
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
by: Zhao, Zengqun, et al.
Published: (2026)
by: Zhao, Zengqun, et al.
Published: (2026)
Video Generation with Predictive Latents
by: Zhao, Yian, et al.
Published: (2026)
by: Zhao, Yian, et al.
Published: (2026)
HairDiffusion: Vivid Multi-Colored Hair Editing via Latent Diffusion
by: Zeng, Yu, et al.
Published: (2024)
by: Zeng, Yu, et al.
Published: (2024)
Similar Items
-
Concat-ID: Towards Universal Identity-Preserving Video Synthesis
by: Zhong, Yong, et al.
Published: (2025) -
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
by: Zheng, Wendi, et al.
Published: (2024) -
Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
by: Zhang, Zhenxing, et al.
Published: (2025) -
SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
by: Yan, Wenhao, et al.
Published: (2025) -
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)