Large Motion Video Autoencoding with Cross-modal Video VAE
Fuente:
arXiv
Saved in:
| Main Authors: | Xing, Yazhou, Fei, Yang, He, Yingqing, Chen, Jingye, Xie, Jiaxin, Chi, Xiaowei, Chen, Qifeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
by: Pham, Kien T., et al.
Published: (2025)
by: Pham, Kien T., et al.
Published: (2025)
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
by: Rao, Zhefan, et al.
Published: (2024)
by: Rao, Zhefan, et al.
Published: (2024)
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
by: Fang, Pengjun, et al.
Published: (2026)
by: Fang, Pengjun, et al.
Published: (2026)
AnimationBench: Are Video Models Good at Character-Centric Animation?
by: Wu, Leyi, et al.
Published: (2026)
by: Wu, Leyi, et al.
Published: (2026)
MagicStick: Controllable Video Editing via Control Handle Transformations
by: Ma, Yue, et al.
Published: (2023)
by: Ma, Yue, et al.
Published: (2023)
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
by: Xing, Yazhou, et al.
Published: (2024)
by: Xing, Yazhou, et al.
Published: (2024)
TALE: Training-free Cross-domain Image Composition via Adaptive Latent Manipulation and Energy-guided Optimization
by: Pham, Kien T., et al.
Published: (2024)
by: Pham, Kien T., et al.
Published: (2024)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
by: Ma, Yue, et al.
Published: (2023)
by: Ma, Yue, et al.
Published: (2023)
VideoMAC: Video Masked Autoencoders Meet ConvNets
by: Pei, Gensheng, et al.
Published: (2024)
by: Pei, Gensheng, et al.
Published: (2024)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022)
by: Li, Bing, et al.
Published: (2022)
Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
by: Fei, Yang, et al.
Published: (2025)
by: Fei, Yang, et al.
Published: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
LLMs Meet Multimodal Generation and Editing: A Survey
by: He, Yingqing, et al.
Published: (2024)
by: He, Yingqing, et al.
Published: (2024)
LDM-ISP: Enhancing Neural ISP for Low Light with Latent Diffusion Models
by: Wen, Qiang, et al.
Published: (2023)
by: Wen, Qiang, et al.
Published: (2023)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
VideoXum: Cross-modal Visual and Textural Summarization of Videos
by: Lin, Jingyang, et al.
Published: (2023)
by: Lin, Jingyang, et al.
Published: (2023)
Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization
by: Zhu, Zhiyi, et al.
Published: (2025)
by: Zhu, Zhiyi, et al.
Published: (2025)
FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
by: Qiu, Haonan, et al.
Published: (2024)
by: Qiu, Haonan, et al.
Published: (2024)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
by: Zhou, Yupeng, et al.
Published: (2025)
by: Zhou, Yupeng, et al.
Published: (2025)
DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation
by: Yuan, Zhihang, et al.
Published: (2025)
by: Yuan, Zhihang, et al.
Published: (2025)
Cross-modal Causal Relation Alignment for Video Question Grounding
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
Does Synthetic Layered Design Data Benefit Layered Design Decomposition?
by: Wu, Kam Man, et al.
Published: (2026)
by: Wu, Kam Man, et al.
Published: (2026)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
by: Nie, Yuxiang, et al.
Published: (2025)
by: Nie, Yuxiang, et al.
Published: (2025)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
by: Chen, Liuhan, et al.
Published: (2024)
by: Chen, Liuhan, et al.
Published: (2024)
Asymmetric VAE for One-Step Video Super-Resolution Acceleration
by: Li, Jianze, et al.
Published: (2025)
by: Li, Jianze, et al.
Published: (2025)
CoMo: Compositional Motion Customization for Text-to-Video Generation
by: Xu, Youcan, et al.
Published: (2025)
by: Xu, Youcan, et al.
Published: (2025)
FastVMT: Eliminating Redundancy in Video Motion Transfer
by: Ma, Yue, et al.
Published: (2026)
by: Ma, Yue, et al.
Published: (2026)
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
by: Ahamed, Shihab Aaqil, et al.
Published: (2025)
by: Ahamed, Shihab Aaqil, et al.
Published: (2025)
Improved Video VAE for Latent Video Diffusion Model
by: Wu, Pingyu, et al.
Published: (2024)
by: Wu, Pingyu, et al.
Published: (2024)
Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
VidTwin: Video VAE with Decoupled Structure and Dynamics
by: Wang, Yuchi, et al.
Published: (2024)
by: Wang, Yuchi, et al.
Published: (2024)
Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation
by: Liang, Chen, et al.
Published: (2021)
by: Liang, Chen, et al.
Published: (2021)
Recurrent Video Masked Autoencoders
by: Zoran, Daniel, et al.
Published: (2025)
by: Zoran, Daniel, et al.
Published: (2025)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
Similar Items
-
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
by: Pham, Kien T., et al.
Published: (2025) -
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
by: Liu, Huaize, et al.
Published: (2025) -
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
by: Rao, Zhefan, et al.
Published: (2024) -
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
by: Fang, Pengjun, et al.
Published: (2026) -
AnimationBench: Are Video Models Good at Character-Centric Animation?
by: Wu, Leyi, et al.
Published: (2026)