Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Huaize, Sun, Wenzhang, Zhang, Qiyuan, Di, Donglin, Gong, Biao, Li, Hao, Wei, Chen, Zou, Changqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Self-supervised Motion Representation for Portrait Video Generation
by: Zhang, Qiyuan, et al.
Published: (2025)
by: Zhang, Qiyuan, et al.
Published: (2025)
MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control
by: Sun, Wenzhang, et al.
Published: (2024)
by: Sun, Wenzhang, et al.
Published: (2024)
Large Motion Video Autoencoding with Cross-modal Video VAE
by: Xing, Yazhou, et al.
Published: (2024)
by: Xing, Yazhou, et al.
Published: (2024)
UniCP: A Unified Caching and Pruning Framework for Efficient Video Generation
by: Sun, Wenzhang, et al.
Published: (2025)
by: Sun, Wenzhang, et al.
Published: (2025)
DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
by: Di, Donglin, et al.
Published: (2024)
by: Di, Donglin, et al.
Published: (2024)
DeCo-VAE: Learning Compact Latents for Video Reconstruction via Decoupled Representation
by: Yin, Xiangchen, et al.
Published: (2025)
by: Yin, Xiangchen, et al.
Published: (2025)
ChronoTailor: Harnessing Attention Guidance for Fine-Grained Video Virtual Try-On
by: Wang, Jinjuan, et al.
Published: (2025)
by: Wang, Jinjuan, et al.
Published: (2025)
PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models
by: Zhang, Qiyuan, et al.
Published: (2026)
by: Zhang, Qiyuan, et al.
Published: (2026)
SeNM-VAE: Semi-Supervised Noise Modeling with Hierarchical Variational Autoencoder
by: Zheng, Dihan, et al.
Published: (2024)
by: Zheng, Dihan, et al.
Published: (2024)
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
by: Zou, Kai, et al.
Published: (2026)
by: Zou, Kai, et al.
Published: (2026)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Eliminating VAE for Fast and High-Resolution Generative Detail Restoration
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
Real Face Video Animation Platform
by: Chen, Xiaokai, et al.
Published: (2024)
by: Chen, Xiaokai, et al.
Published: (2024)
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
by: Tan, Shuai, et al.
Published: (2025)
by: Tan, Shuai, et al.
Published: (2025)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
by: Cai, Xin, et al.
Published: (2026)
by: Cai, Xin, et al.
Published: (2026)
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
by: Zhu, Lunjie, et al.
Published: (2026)
by: Zhu, Lunjie, et al.
Published: (2026)
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
by: Ma, Yongjia, et al.
Published: (2025)
by: Ma, Yongjia, et al.
Published: (2025)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
by: Zhou, Yupeng, et al.
Published: (2025)
by: Zhou, Yupeng, et al.
Published: (2025)
Leveraging SAM for Single-Source Domain Generalization in Medical Image Segmentation
by: Wang, Hanhui, et al.
Published: (2024)
by: Wang, Hanhui, et al.
Published: (2024)
Learning Human Motion from Monocular Videos via Cross-Modal Manifold Alignment
by: Hou, Shuaiying, et al.
Published: (2024)
by: Hou, Shuaiying, et al.
Published: (2024)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
by: Yang, Jiahui, et al.
Published: (2024)
by: Yang, Jiahui, et al.
Published: (2024)
DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders
by: Wang, Tianhang, et al.
Published: (2026)
by: Wang, Tianhang, et al.
Published: (2026)
GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
by: Ling, Zeyu, et al.
Published: (2024)
by: Ling, Zeyu, et al.
Published: (2024)
HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
by: Ding, Xinpeng, et al.
Published: (2023)
by: Ding, Xinpeng, et al.
Published: (2023)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
Hi-ResNet: Edge Detail Enhancement for High-Resolution Remote Sensing Segmentation
by: Chen, Yuxia, et al.
Published: (2023)
by: Chen, Yuxia, et al.
Published: (2023)
Focus-Consistent Multi-Level Aggregation for Compositional Zero-Shot Learning
by: Dai, Fengyuan, et al.
Published: (2024)
by: Dai, Fengyuan, et al.
Published: (2024)
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
by: Shi, Shuwei, et al.
Published: (2024)
by: Shi, Shuwei, et al.
Published: (2024)
VideoMAC: Video Masked Autoencoders Meet ConvNets
by: Pei, Gensheng, et al.
Published: (2024)
by: Pei, Gensheng, et al.
Published: (2024)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration
by: Sun, Wenzhang, et al.
Published: (2026)
by: Sun, Wenzhang, et al.
Published: (2026)
Efficient Motion-Aware Video MLLM
by: Zhao, Zijia, et al.
Published: (2025)
by: Zhao, Zijia, et al.
Published: (2025)
Improved Video VAE for Latent Video Diffusion Model
by: Wu, Pingyu, et al.
Published: (2024)
by: Wu, Pingyu, et al.
Published: (2024)
HiFiVFS: High Fidelity Video Face Swapping
by: Chen, Xu, et al.
Published: (2024)
by: Chen, Xu, et al.
Published: (2024)
Similar Items
-
A Self-supervised Motion Representation for Portrait Video Generation
by: Zhang, Qiyuan, et al.
Published: (2025) -
MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation
by: Liu, Huaize, et al.
Published: (2025) -
UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control
by: Sun, Wenzhang, et al.
Published: (2024) -
Large Motion Video Autoencoding with Cross-modal Video VAE
by: Xing, Yazhou, et al.
Published: (2024) -
UniCP: A Unified Caching and Pruning Framework for Efficient Video Generation
by: Sun, Wenzhang, et al.
Published: (2025)