StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tu, Shuyuan, Pan, Yueming, Huang, Yinming, Han, Xintong, Xing, Zhen, Dai, Qi, Luo, Chong, Wu, Zuxuan, Jiang, Yu-Gang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913984987267072
author Tu, Shuyuan
Pan, Yueming
Huang, Yinming
Han, Xintong
Xing, Zhen
Dai, Qi
Luo, Chong
Wu, Zuxuan
Jiang, Yu-Gang
author_facet Tu, Shuyuan
Pan, Yueming
Huang, Yinming
Han, Xintong
Xing, Zhen
Dai, Qi
Luo, Chong
Wu, Zuxuan
Jiang, Yu-Gang
contents Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end video diffusion transformer that synthesizes infinite-length high-quality videos without post-processing. Conditioned on a reference image and audio, StableAvatar integrates tailored training and inference modules to enable infinite-length video generation. We observe that the main reason preventing existing models from generating long videos lies in their audio modeling. They typically rely on third-party off-the-shelf extractors to obtain audio embeddings, which are then directly injected into the diffusion model via cross-attention. Since current diffusion backbones lack any audio-related priors, this approach causes severe latent distribution error accumulation across video clips, leading the latent distribution of subsequent segments to drift away from the optimal distribution gradually. To address this, StableAvatar introduces a novel Time-step-aware Audio Adapter that prevents error accumulation via time-step-aware modulation. During inference, we propose a novel Audio Native Guidance Mechanism to further enhance the audio synchronization by leveraging the diffusion's own evolving joint audio-latent prediction as a dynamic guidance signal. To enhance the smoothness of the infinite-length videos, we introduce a Dynamic Weighted Sliding-window Strategy that fuses latent over time. Experiments on benchmarks show the effectiveness of StableAvatar both qualitatively and quantitatively.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08248
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
Tu, Shuyuan
Pan, Yueming
Huang, Yinming
Han, Xintong
Xing, Zhen
Dai, Qi
Luo, Chong
Wu, Zuxuan
Jiang, Yu-Gang
Computer Vision and Pattern Recognition
Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end video diffusion transformer that synthesizes infinite-length high-quality videos without post-processing. Conditioned on a reference image and audio, StableAvatar integrates tailored training and inference modules to enable infinite-length video generation. We observe that the main reason preventing existing models from generating long videos lies in their audio modeling. They typically rely on third-party off-the-shelf extractors to obtain audio embeddings, which are then directly injected into the diffusion model via cross-attention. Since current diffusion backbones lack any audio-related priors, this approach causes severe latent distribution error accumulation across video clips, leading the latent distribution of subsequent segments to drift away from the optimal distribution gradually. To address this, StableAvatar introduces a novel Time-step-aware Audio Adapter that prevents error accumulation via time-step-aware modulation. During inference, we propose a novel Audio Native Guidance Mechanism to further enhance the audio synchronization by leveraging the diffusion's own evolving joint audio-latent prediction as a dynamic guidance signal. To enhance the smoothness of the infinite-length videos, we introduce a Dynamic Weighted Sliding-window Strategy that fuses latent over time. Experiments on benchmarks show the effectiveness of StableAvatar both qualitatively and quantitatively.
title StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.08248