OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yupeng, Li, Zhen, Ouyang, Ziheng, Chen, Yuming, Du, Ruoyi, Zhou, Daquan, Fu, Bin, Liu, Yihao, Gao, Peng, Cheng, Ming-Ming, Hou, Qibin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024)
by: Zhou, Yupeng, et al.
Published: (2024)
K-LoRA: Unlocking Training-Free Fusion of Any Subject and Style LoRAs
by: Ouyang, Ziheng, et al.
Published: (2025)
by: Ouyang, Ziheng, et al.
Published: (2025)
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
by: Zhou, Yupeng, et al.
Published: (2026)
by: Zhou, Yupeng, et al.
Published: (2026)
Sora Generates Videos with Stunning Geometrical Consistency
by: Li, Xuanyi, et al.
Published: (2024)
by: Li, Xuanyi, et al.
Published: (2024)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
by: Zhang, Xuying, et al.
Published: (2025)
by: Zhang, Xuying, et al.
Published: (2025)
HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
by: Takida, Yuhta, et al.
Published: (2023)
by: Takida, Yuhta, et al.
Published: (2023)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
EdVAE: Mitigating Codebook Collapse with Evidential Discrete Variational Autoencoders
by: Baykal, Gulcin, et al.
Published: (2023)
by: Baykal, Gulcin, et al.
Published: (2023)
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
Quantize-then-Rectify: Efficient VQ-VAE Training
by: Zhang, Borui, et al.
Published: (2025)
by: Zhang, Borui, et al.
Published: (2025)
LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation
by: Yuan, Zhihang, et al.
Published: (2025)
by: Yuan, Zhihang, et al.
Published: (2025)
Asymmetric VAE for One-Step Video Super-Resolution Acceleration
by: Li, Jianze, et al.
Published: (2025)
by: Li, Jianze, et al.
Published: (2025)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
by: Chen, Liuhan, et al.
Published: (2024)
by: Chen, Liuhan, et al.
Published: (2024)
High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
by: Zeng, Quan-Sheng, et al.
Published: (2024)
by: Zeng, Quan-Sheng, et al.
Published: (2024)
SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
by: Zhou, Yupeng, et al.
Published: (2023)
by: Zhou, Yupeng, et al.
Published: (2023)
HumanNet: Scaling Human-centric Video Learning to One Million Hours
by: Deng, Yufan, et al.
Published: (2026)
by: Deng, Yufan, et al.
Published: (2026)
Purification for Hybrid Entanglement between Discrete‐ and Continuous‐Variable States
by: Cheng‐Chen Luo, et al.
Published: (2024)
by: Cheng‐Chen Luo, et al.
Published: (2024)
Improved Video VAE for Latent Video Diffusion Model
by: Wu, Pingyu, et al.
Published: (2024)
by: Wu, Pingyu, et al.
Published: (2024)
CoVAE: Consistency Training of Variational Autoencoders
by: Silvestri, Gianluigi, et al.
Published: (2025)
by: Silvestri, Gianluigi, et al.
Published: (2025)
Mixture of Style Experts for Diverse Image Stylization
by: Zhu, Shihao, et al.
Published: (2026)
by: Zhu, Shihao, et al.
Published: (2026)
Half-VAE: An Encoder-Free VAE to Bypass Explicit Inverse Mapping
by: Wei, Yuan-Hao, et al.
Published: (2024)
by: Wei, Yuan-Hao, et al.
Published: (2024)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
by: Niu, Zhikang, et al.
Published: (2025)
by: Niu, Zhikang, et al.
Published: (2025)
Attentive VQ-VAE
by: Hoyos, Angello, et al.
Published: (2023)
by: Hoyos, Angello, et al.
Published: (2023)
Large Motion Video Autoencoding with Cross-modal Video VAE
by: Xing, Yazhou, et al.
Published: (2024)
by: Xing, Yazhou, et al.
Published: (2024)
GQ-VAE: A gated quantized VAE for learning variable length tokens
by: Datta, Theo, et al.
Published: (2025)
by: Datta, Theo, et al.
Published: (2025)
Support is All You Need for Certified VAE Training
by: Xu, Changming, et al.
Published: (2025)
by: Xu, Changming, et al.
Published: (2025)
VidTwin: Video VAE with Decoupled Structure and Dynamics
by: Wang, Yuchi, et al.
Published: (2024)
by: Wang, Yuchi, et al.
Published: (2024)
Comparative Analysis of MDL-VAE vs. Standard VAE on 202 Years of Gynecological Data
by: Santos, Paula
Published: (2025)
by: Santos, Paula
Published: (2025)
HireVAE: An Online and Adaptive Factor Model Based on Hierarchical and Regime-Switch VAE
by: Wei, Zikai, et al.
Published: (2023)
by: Wei, Zikai, et al.
Published: (2023)
SepVAE: a contrastive VAE to separate pathological patterns from healthy ones
by: Louiset, Robin, et al.
Published: (2023)
by: Louiset, Robin, et al.
Published: (2023)
Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
Time-Causal VAE: Robust Financial Time Series Generator
by: Acciaio, Beatrice, et al.
Published: (2024)
by: Acciaio, Beatrice, et al.
Published: (2024)
Disentangle VAE for Molecular Generation
by: Wang, Yanbo, et al.
Published: (2022)
by: Wang, Yanbo, et al.
Published: (2022)
Iterative Amortized Hierarchical VAE
by: Penninga, Simon W., et al.
Published: (2026)
by: Penninga, Simon W., et al.
Published: (2026)
How to train your VAE
by: Rivera, Mariano
Published: (2023)
by: Rivera, Mariano
Published: (2023)
VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
by: Novitskiy, Lev, et al.
Published: (2025)
by: Novitskiy, Lev, et al.
Published: (2025)
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
by: Ouyang, Ziheng, et al.
Published: (2025)
by: Ouyang, Ziheng, et al.
Published: (2025)
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
Similar Items
-
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024) -
K-LoRA: Unlocking Training-Free Fusion of Any Subject and Style LoRAs
by: Ouyang, Ziheng, et al.
Published: (2025) -
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
by: Zhou, Yupeng, et al.
Published: (2026) -
Sora Generates Videos with Stunning Geometrical Consistency
by: Li, Xuanyi, et al.
Published: (2024) -
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
by: Zhang, Xuying, et al.
Published: (2025)