VidTwin: Video VAE with Decoupled Structure and Dynamics

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Yuchi, Guo, Junliang, Xie, Xinyi, He, Tianyu, Sun, Xu, Bian, Jiang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915217817993216
author Wang, Yuchi
Guo, Junliang
Xie, Xinyi
He, Tianyu
Sun, Xu
Bian, Jiang
author_facet Wang, Yuchi
Guo, Junliang
Xie, Xinyi
He, Tianyu
Sun, Xu
Bian, Jiang
contents Recent advancements in video autoencoders (Video AEs) have significantly improved the quality and efficiency of video generation. In this paper, we propose a novel and compact video autoencoder, VidTwin, that decouples video into two distinct latent spaces: Structure latent vectors, which capture overall content and global movement, and Dynamics latent vectors, which represent fine-grained details and rapid movements. Specifically, our approach leverages an Encoder-Decoder backbone, augmented with two submodules for extracting these latent spaces, respectively. The first submodule employs a Q-Former to extract low-frequency motion trends, followed by downsampling blocks to remove redundant content details. The second averages the latent vectors along the spatial dimension to capture rapid motion. Extensive experiments show that VidTwin achieves a high compression rate of 0.20% with high reconstruction quality (PSNR of 28.14 on the MCL-JCV dataset), and performs efficiently and effectively in downstream generative tasks. Moreover, our model demonstrates explainability and scalability, paving the way for future research in video latent representation and generation. Check our project page for more details: https://vidtwin.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17726
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VidTwin: Video VAE with Decoupled Structure and Dynamics
Wang, Yuchi
Guo, Junliang
Xie, Xinyi
He, Tianyu
Sun, Xu
Bian, Jiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent advancements in video autoencoders (Video AEs) have significantly improved the quality and efficiency of video generation. In this paper, we propose a novel and compact video autoencoder, VidTwin, that decouples video into two distinct latent spaces: Structure latent vectors, which capture overall content and global movement, and Dynamics latent vectors, which represent fine-grained details and rapid movements. Specifically, our approach leverages an Encoder-Decoder backbone, augmented with two submodules for extracting these latent spaces, respectively. The first submodule employs a Q-Former to extract low-frequency motion trends, followed by downsampling blocks to remove redundant content details. The second averages the latent vectors along the spatial dimension to capture rapid motion. Extensive experiments show that VidTwin achieves a high compression rate of 0.20% with high reconstruction quality (PSNR of 28.14 on the MCL-JCV dataset), and performs efficiently and effectively in downstream generative tasks. Moreover, our model demonstrates explainability and scalability, paving the way for future research in video latent representation and generation. Check our project page for more details: https://vidtwin.github.io/.
title VidTwin: Video VAE with Decoupled Structure and Dynamics
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.17726