Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Lunjie, Huang, Yushi, Ge, Xingtong, Xue, Yufei, Liu, Zhening, Zhang, Yumeng, Lin, Zehong, Zhang, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Real-Time Human Frontal View Synthesis from a Single Image
by: Lin, Fangyu, et al.
Published: (2026)
by: Lin, Fangyu, et al.
Published: (2026)
LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
by: Huang, Yushi, et al.
Published: (2025)
by: Huang, Yushi, et al.
Published: (2025)
MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes
by: Zhang, Xinjie, et al.
Published: (2024)
by: Zhang, Xinjie, et al.
Published: (2024)
VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
by: Xue, Yufei, et al.
Published: (2025)
by: Xue, Yufei, et al.
Published: (2025)
Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling
by: Liu, Zhening, et al.
Published: (2025)
by: Liu, Zhening, et al.
Published: (2025)
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
by: Zou, Ya, et al.
Published: (2025)
by: Zou, Ya, et al.
Published: (2025)
EVA-Gaussian: 3D Gaussian-based Real-time Human Novel View Synthesis under Diverse Multi-view Camera Settings
by: Hu, Yingdong, et al.
Published: (2024)
by: Hu, Yingdong, et al.
Published: (2024)
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
by: Zhang, Yiming, et al.
Published: (2023)
by: Zhang, Yiming, et al.
Published: (2023)
Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
by: Ge, Xingtong, et al.
Published: (2026)
by: Ge, Xingtong, et al.
Published: (2026)
Bidirectional Stereo Image Compression with Cross-Dimensional Entropy Model
by: Liu, Zhening, et al.
Published: (2024)
by: Liu, Zhening, et al.
Published: (2024)
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
by: Cai, Xin, et al.
Published: (2026)
by: Cai, Xin, et al.
Published: (2026)
RemedyGS: Defend 3D Gaussian Splatting against Computation Cost Attacks
by: Li, Yanping, et al.
Published: (2025)
by: Li, Yanping, et al.
Published: (2025)
Mon3tr: Monocular 3D Telepresence with Pre-built Gaussian Avatars as Amortization
by: Lin, Fangyu, et al.
Published: (2026)
by: Lin, Fangyu, et al.
Published: (2026)
Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
by: Xue, Bowen, et al.
Published: (2025)
by: Xue, Bowen, et al.
Published: (2025)
Boosting Neural Representations for Videos with a Conditional Decoder
by: Zhang, Xinjie, et al.
Published: (2024)
by: Zhang, Xinjie, et al.
Published: (2024)
Dynamics-Aware Gaussian Splatting Streaming Towards Fast On-the-Fly 4D Reconstruction
by: Liu, Zhening, et al.
Published: (2024)
by: Liu, Zhening, et al.
Published: (2024)
FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation
by: Zhang, Liuzhou, et al.
Published: (2026)
by: Zhang, Liuzhou, et al.
Published: (2026)
Plug-and-Play Versatile Compressed Video Enhancement
by: Zeng, Huimin, et al.
Published: (2025)
by: Zeng, Huimin, et al.
Published: (2025)
Plug-and-Play Context Feature Reuse for Efficient Masked Generation
by: Liu, Xuejie, et al.
Published: (2025)
by: Liu, Xuejie, et al.
Published: (2025)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
by: Liu, Xuyang, et al.
Published: (2025)
by: Liu, Xuyang, et al.
Published: (2025)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
by: Lv, Chengtao, et al.
Published: (2025)
by: Lv, Chengtao, et al.
Published: (2025)
Plug-and-Play Diffusion Distillation
by: Hsiao, Yi-Ting, et al.
Published: (2024)
by: Hsiao, Yi-Ting, et al.
Published: (2024)
Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning
by: Yan, Zhiyuan, et al.
Published: (2024)
by: Yan, Zhiyuan, et al.
Published: (2024)
PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation
by: Chen, Yongwei, et al.
Published: (2026)
by: Chen, Yongwei, et al.
Published: (2026)
Score-Based Turbo Message Passing for Plug-and-Play Compressive Imaging
by: Cai, Chang, et al.
Published: (2025)
by: Cai, Chang, et al.
Published: (2025)
DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation
by: Yuan, Zhihang, et al.
Published: (2025)
by: Yuan, Zhihang, et al.
Published: (2025)
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
by: Ma, Bingqi, et al.
Published: (2026)
by: Ma, Bingqi, et al.
Published: (2026)
CBNet: A Plug-and-Play Network for Segmentation-Based Scene Text Detection
by: Zhao, Xi, et al.
Published: (2022)
by: Zhao, Xi, et al.
Published: (2022)
GI-GS: Global Illumination Decomposition on Gaussian Splatting for Inverse Rendering
by: Chen, Hongze, et al.
Published: (2024)
by: Chen, Hongze, et al.
Published: (2024)
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
by: Zhang, Xinjie, et al.
Published: (2024)
by: Zhang, Xinjie, et al.
Published: (2024)
Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
by: Lv, Chengtao, et al.
Published: (2026)
by: Lv, Chengtao, et al.
Published: (2026)
Text-to-Image Rectified Flow as Plug-and-Play Priors
by: Yang, Xiaofeng, et al.
Published: (2024)
by: Yang, Xiaofeng, et al.
Published: (2024)
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
by: He, Dailan, et al.
Published: (2026)
by: He, Dailan, et al.
Published: (2026)
ReMA: A Training-Free Plug-and-Play Mixing Augmentation for Video Behavior Recognition
by: Cui, Feng-Qi, et al.
Published: (2026)
by: Cui, Feng-Qi, et al.
Published: (2026)
Spatial-Aware Latent Initialization for Controllable Image Generation
by: Sun, Wenqiang, et al.
Published: (2024)
by: Sun, Wenqiang, et al.
Published: (2024)
Similar Items
-
Real-Time Human Frontal View Synthesis from a Single Image
by: Lin, Fangyu, et al.
Published: (2026) -
LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
by: Huang, Yushi, et al.
Published: (2025) -
MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes
by: Zhang, Xinjie, et al.
Published: (2024) -
VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
by: Xue, Yufei, et al.
Published: (2025) -
Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling
by: Liu, Zhening, et al.
Published: (2025)