InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Yuchi, Guo, Junliang, Bai, Jianhong, Yu, Runyi, He, Tianyu, Tan, Xu, Sun, Xu, Bian, Jiang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929357629423616
author Wang, Yuchi
Guo, Junliang
Bai, Jianhong
Yu, Runyi
He, Tianyu
Tan, Xu
Sun, Xu
Bian, Jiang
author_facet Wang, Yuchi
Guo, Junliang
Bai, Jianhong
Yu, Runyi
He, Tianyu
Tan, Xu
Sun, Xu
Bian, Jiang
contents Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the generated video less vivid and controllable. In this paper, we propose a novel text-guided approach for generating emotionally expressive 2D avatars, offering fine-grained control, improved interactivity, and generalizability to the resulting video. Our framework, named InstructAvatar, leverages a natural language interface to control the emotion as well as the facial motion of avatars. Technically, we design an automatic annotation pipeline to construct an instruction-video paired training dataset, equipped with a novel two-branch diffusion-based generator to predict avatars with audio and text instructions at the same time. Experimental results demonstrate that InstructAvatar produces results that align well with both conditions, and outperforms existing methods in fine-grained emotion control, lip-sync quality, and naturalness. Our project page is https://wangyuchi369.github.io/InstructAvatar/.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15758
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
Wang, Yuchi
Guo, Junliang
Bai, Jianhong
Yu, Runyi
He, Tianyu
Tan, Xu
Sun, Xu
Bian, Jiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the generated video less vivid and controllable. In this paper, we propose a novel text-guided approach for generating emotionally expressive 2D avatars, offering fine-grained control, improved interactivity, and generalizability to the resulting video. Our framework, named InstructAvatar, leverages a natural language interface to control the emotion as well as the facial motion of avatars. Technically, we design an automatic annotation pipeline to construct an instruction-video paired training dataset, equipped with a novel two-branch diffusion-based generator to predict avatars with audio and text instructions at the same time. Experimental results demonstrate that InstructAvatar produces results that align well with both conditions, and outperforms existing methods in fine-grained emotion control, lip-sync quality, and naturalness. Our project page is https://wangyuchi369.github.io/InstructAvatar/.
title InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2405.15758