Wan-S2V: Audio-Driven Cinematic Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Xin, Hu, Li, Hu, Siqi, Huang, Mingyang, Ji, Chaonan, Meng, Dechao, Qi, Jinwei, Qiao, Penchong, Shen, Zhen, Song, Yafei, Sun, Ke, Tian, Linrui, Wang, Guangyuan, Wang, Qi, Wang, Zhongjian, Xiao, Jiayu, Xu, Sheng, Zhang, Bang, Zhang, Peng, Zhang, Xindi, Zhang, Zhe, Zhou, Jingren, Zhuo, Lian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914005779480576
author Gao, Xin
Hu, Li
Hu, Siqi
Huang, Mingyang
Ji, Chaonan
Meng, Dechao
Qi, Jinwei
Qiao, Penchong
Shen, Zhen
Song, Yafei
Sun, Ke
Tian, Linrui
Wang, Guangyuan
Wang, Qi
Wang, Zhongjian
Xiao, Jiayu
Xu, Sheng
Zhang, Bang
Zhang, Peng
Zhang, Xindi
Zhang, Zhe
Zhou, Jingren
Zhuo, Lian
author_facet Gao, Xin
Hu, Li
Hu, Siqi
Huang, Mingyang
Ji, Chaonan
Meng, Dechao
Qi, Jinwei
Qiao, Penchong
Shen, Zhen
Song, Yafei
Sun, Ke
Tian, Linrui
Wang, Guangyuan
Wang, Qi
Wang, Zhongjian
Xiao, Jiayu
Xu, Sheng
Zhang, Bang
Zhang, Peng
Zhang, Xindi
Zhang, Zhe
Zhou, Jingren
Zhuo, Lian
contents Current state-of-the-art (SOTA) methods for audio-driven character animation demonstrate promising performance for scenarios primarily involving speech and singing. However, they often fall short in more complex film and television productions, which demand sophisticated elements such as nuanced character interactions, realistic body movements, and dynamic camera work. To address this long-standing challenge of achieving film-level character animation, we propose an audio-driven model, which we refere to as Wan-S2V, built upon Wan. Our model achieves significantly enhanced expressiveness and fidelity in cinematic contexts compared to existing approaches. We conducted extensive experiments, benchmarking our method against cutting-edge models such as Hunyuan-Avatar and Omnihuman. The experimental results consistently demonstrate that our approach significantly outperforms these existing solutions. Additionally, we explore the versatility of our method through its applications in long-form video generation and precise video lip-sync editing.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Wan-S2V: Audio-Driven Cinematic Video Generation
Gao, Xin
Hu, Li
Hu, Siqi
Huang, Mingyang
Ji, Chaonan
Meng, Dechao
Qi, Jinwei
Qiao, Penchong
Shen, Zhen
Song, Yafei
Sun, Ke
Tian, Linrui
Wang, Guangyuan
Wang, Qi
Wang, Zhongjian
Xiao, Jiayu
Xu, Sheng
Zhang, Bang
Zhang, Peng
Zhang, Xindi
Zhang, Zhe
Zhou, Jingren
Zhuo, Lian
Computer Vision and Pattern Recognition
Current state-of-the-art (SOTA) methods for audio-driven character animation demonstrate promising performance for scenarios primarily involving speech and singing. However, they often fall short in more complex film and television productions, which demand sophisticated elements such as nuanced character interactions, realistic body movements, and dynamic camera work. To address this long-standing challenge of achieving film-level character animation, we propose an audio-driven model, which we refere to as Wan-S2V, built upon Wan. Our model achieves significantly enhanced expressiveness and fidelity in cinematic contexts compared to existing approaches. We conducted extensive experiments, benchmarking our method against cutting-edge models such as Hunyuan-Avatar and Omnihuman. The experimental results consistently demonstrate that our approach significantly outperforms these existing solutions. Additionally, we explore the versatility of our method through its applications in long-form video generation and precise video lip-sync editing.
title Wan-S2V: Audio-Driven Cinematic Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.18621