Controllable Video Generation with Provable Disentanglement

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shen, Yifan, Zhu, Peiyuan, Li, Zijian, Xie, Shaoan, Deka, Namrata, Liu, Zongfang, Tang, Zeyu, Chen, Guangyi, Zhang, Kun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912627939082240
author Shen, Yifan
Zhu, Peiyuan
Li, Zijian
Xie, Shaoan
Deka, Namrata
Liu, Zongfang
Tang, Zeyu
Chen, Guangyi
Zhang, Kun
author_facet Shen, Yifan
Zhu, Peiyuan
Li, Zijian
Xie, Shaoan
Deka, Namrata
Liu, Zongfang
Tang, Zeyu
Chen, Guangyi
Zhang, Kun
contents Controllable video generation remains a significant challenge, despite recent advances in generating high-quality and consistent videos. Most existing methods for controlling video generation treat the video as a whole, neglecting intricate fine-grained spatiotemporal relationships, which limits both control precision and efficiency. In this paper, we propose Controllable Video Generative Adversarial Networks (CoVoGAN) to disentangle the video concepts, thus facilitating efficient and independent control over individual concepts. Specifically, following the minimal change principle, we first disentangle static and dynamic latent variables. We then leverage the sufficient change property to achieve component-wise identifiability of dynamic latent variables, enabling disentangled control of video generation. To establish the theoretical foundation, we provide a rigorous analysis demonstrating the identifiability of our approach. Building on these theoretical insights, we design a Temporal Transition Module to disentangle latent dynamics. To enforce the minimal change principle and sufficient change property, we minimize the dimensionality of latent dynamic variables and impose temporal conditional independence. To validate our approach, we integrate this module as a plug-in for GANs. Extensive qualitative and quantitative experiments on various video generation benchmarks demonstrate that our method significantly improves generation quality and controllability across diverse real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02690
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Controllable Video Generation with Provable Disentanglement
Shen, Yifan
Zhu, Peiyuan
Li, Zijian
Xie, Shaoan
Deka, Namrata
Liu, Zongfang
Tang, Zeyu
Chen, Guangyi
Zhang, Kun
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Controllable video generation remains a significant challenge, despite recent advances in generating high-quality and consistent videos. Most existing methods for controlling video generation treat the video as a whole, neglecting intricate fine-grained spatiotemporal relationships, which limits both control precision and efficiency. In this paper, we propose Controllable Video Generative Adversarial Networks (CoVoGAN) to disentangle the video concepts, thus facilitating efficient and independent control over individual concepts. Specifically, following the minimal change principle, we first disentangle static and dynamic latent variables. We then leverage the sufficient change property to achieve component-wise identifiability of dynamic latent variables, enabling disentangled control of video generation. To establish the theoretical foundation, we provide a rigorous analysis demonstrating the identifiability of our approach. Building on these theoretical insights, we design a Temporal Transition Module to disentangle latent dynamics. To enforce the minimal change principle and sufficient change property, we minimize the dimensionality of latent dynamic variables and impose temporal conditional independence. To validate our approach, we integrate this module as a plug-in for GANs. Extensive qualitative and quantitative experiments on various video generation benchmarks demonstrate that our method significantly improves generation quality and controllability across diverse real-world scenarios.
title Controllable Video Generation with Provable Disentanglement
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.02690