OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gan, Qijun, Yang, Ruizi, Zhu, Jianke, Xue, Shaofei, Hoi, Steven
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912445095739392
author Gan, Qijun
Yang, Ruizi
Zhu, Jianke
Xue, Shaofei
Hoi, Steven
author_facet Gan, Qijun
Yang, Ruizi
Zhu, Jianke
Xue, Shaofei
Hoi, Steven
contents Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and fluidity. They also struggle with precise prompt control for fine-grained generation. To tackle these challenges, we introduce OmniAvatar, an innovative audio-driven full-body video generation model that enhances human animation with improved lip-sync accuracy and natural movements. OmniAvatar introduces a pixel-wise multi-hierarchical audio embedding strategy to better capture audio features in the latent space, enhancing lip-syncing across diverse scenes. To preserve the capability for prompt-driven control of foundation models while effectively incorporating audio features, we employ a LoRA-based training approach. Extensive experiments show that OmniAvatar surpasses existing models in both facial and semi-body video generation, offering precise text-based control for creating videos in various domains, such as podcasts, human interactions, dynamic scenes, and singing. Our project page is https://omni-avatar.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18866
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
Gan, Qijun
Yang, Ruizi
Zhu, Jianke
Xue, Shaofei
Hoi, Steven
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and fluidity. They also struggle with precise prompt control for fine-grained generation. To tackle these challenges, we introduce OmniAvatar, an innovative audio-driven full-body video generation model that enhances human animation with improved lip-sync accuracy and natural movements. OmniAvatar introduces a pixel-wise multi-hierarchical audio embedding strategy to better capture audio features in the latent space, enhancing lip-syncing across diverse scenes. To preserve the capability for prompt-driven control of foundation models while effectively incorporating audio features, we employ a LoRA-based training approach. Extensive experiments show that OmniAvatar surpasses existing models in both facial and semi-body video generation, offering precise text-based control for creating videos in various domains, such as podcasts, human interactions, dynamic scenes, and singing. Our project page is https://omni-avatar.github.io/.
title OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2506.18866