KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Xingrui, Liu, Jiang, Wang, Ze, Yu, Xiaodong, Wu, Jialian, Sun, Ximeng, Su, Yusheng, Yuille, Alan, Liu, Zicheng, Barsoum, Emad
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911212741066752
author Wang, Xingrui
Liu, Jiang
Wang, Ze
Yu, Xiaodong
Wu, Jialian
Sun, Ximeng
Su, Yusheng
Yuille, Alan
Liu, Zicheng
Barsoum, Emad
author_facet Wang, Xingrui
Liu, Jiang
Wang, Ze
Yu, Xiaodong
Wu, Jialian
Sun, Ximeng
Su, Yusheng
Yuille, Alan
Liu, Zicheng
Barsoum, Emad
contents Generating video from various conditions, such as text, image, and audio, enables both spatial and temporal control, leading to high-quality generation results. Videos with dramatic motions often require a higher frame rate to ensure smooth motion. Currently, most audio-to-visual animation models use uniformly sampled frames from video clips. However, these uniformly sampled frames fail to capture significant key moments in dramatic motions at low frame rates and require significantly more memory when increasing the number of frames directly. In this paper, we propose KeyVID, a keyframe-aware audio-to-visual animation framework that significantly improves the generation quality for key moments in audio signals while maintaining computation efficiency. Given an image and an audio input, we first localize keyframe time steps from the audio. Then, we use a keyframe generator to generate the corresponding visual keyframes. Finally, we generate all intermediate frames using the motion interpolator. Through extensive experiments, we demonstrate that KeyVID significantly improves audio-video synchronization and video quality across multiple datasets, particularly for highly dynamic motions. The code is released in https://github.com/XingruiWang/KeyVID.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09656
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
Wang, Xingrui
Liu, Jiang
Wang, Ze
Yu, Xiaodong
Wu, Jialian
Sun, Ximeng
Su, Yusheng
Yuille, Alan
Liu, Zicheng
Barsoum, Emad
Computer Vision and Pattern Recognition
Generating video from various conditions, such as text, image, and audio, enables both spatial and temporal control, leading to high-quality generation results. Videos with dramatic motions often require a higher frame rate to ensure smooth motion. Currently, most audio-to-visual animation models use uniformly sampled frames from video clips. However, these uniformly sampled frames fail to capture significant key moments in dramatic motions at low frame rates and require significantly more memory when increasing the number of frames directly. In this paper, we propose KeyVID, a keyframe-aware audio-to-visual animation framework that significantly improves the generation quality for key moments in audio signals while maintaining computation efficiency. Given an image and an audio input, we first localize keyframe time steps from the audio. Then, we use a keyframe generator to generate the corresponding visual keyframes. Finally, we generate all intermediate frames using the motion interpolator. Through extensive experiments, we demonstrate that KeyVID significantly improves audio-video synchronization and video quality across multiple datasets, particularly for highly dynamic motions. The code is released in https://github.com/XingruiWang/KeyVID.
title KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.09656