ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Peng, Ziqiao, Chen, Yi, Ma, Yifeng, Zhang, Guozhen, Sun, Zhiyao, Zhou, Zixiang, Zhang, Youliang, Zhou, Zhengguang, Fan, Zhaoxin, Liu, Hongyan, Zhou, Yuan, Lu, Qinglin, He, Jun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909995149295616
author Peng, Ziqiao
Chen, Yi
Ma, Yifeng
Zhang, Guozhen
Sun, Zhiyao
Zhou, Zixiang
Zhang, Youliang
Zhou, Zhengguang
Fan, Zhaoxin
Liu, Hongyan
Zhou, Yuan
Lu, Qinglin
He, Jun
author_facet Peng, Ziqiao
Chen, Yi
Ma, Yifeng
Zhang, Guozhen
Sun, Zhiyao
Zhou, Zixiang
Zhang, Youliang
Zhou, Zhengguang
Fan, Zhaoxin
Liu, Hongyan
Zhou, Yuan
Lu, Qinglin
He, Jun
contents Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency on additional control signals such as pose skeletons. We present ActAvatar, a framework that achieves phase-level precision in action control through textual guidance by capturing both action semantics and temporal context. Our approach introduces three core innovations: (1) Phase-Aware Cross-Attention (PACA), which decomposes prompts into a global base block and temporally-anchored phase blocks, enabling the model to concentrate on phase-relevant tokens for precise temporal-semantic alignment; (2) Progressive Audio-Visual Alignment, which aligns modality influence with the hierarchical feature learning process-early layers prioritize text for establishing action structure while deeper layers emphasize audio for refining lip movements, preventing modality interference; (3) A two-stage training strategy that first establishes robust audio-visual correspondence on diverse data, then injects action control through fine-tuning on structured annotations, maintaining both audio-visual alignment and the model's text-following capabilities. Extensive experiments demonstrate that ActAvatar significantly outperforms state-of-the-art methods in both action control and visual quality.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19546
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
Peng, Ziqiao
Chen, Yi
Ma, Yifeng
Zhang, Guozhen
Sun, Zhiyao
Zhou, Zixiang
Zhang, Youliang
Zhou, Zhengguang
Fan, Zhaoxin
Liu, Hongyan
Zhou, Yuan
Lu, Qinglin
He, Jun
Computer Vision and Pattern Recognition
Despite significant advances in talking avatar generation, existing methods face critical challenges: insufficient text-following capability for diverse actions, lack of temporal alignment between actions and audio content, and dependency on additional control signals such as pose skeletons. We present ActAvatar, a framework that achieves phase-level precision in action control through textual guidance by capturing both action semantics and temporal context. Our approach introduces three core innovations: (1) Phase-Aware Cross-Attention (PACA), which decomposes prompts into a global base block and temporally-anchored phase blocks, enabling the model to concentrate on phase-relevant tokens for precise temporal-semantic alignment; (2) Progressive Audio-Visual Alignment, which aligns modality influence with the hierarchical feature learning process-early layers prioritize text for establishing action structure while deeper layers emphasize audio for refining lip movements, preventing modality interference; (3) A two-stage training strategy that first establishes robust audio-visual correspondence on diverse data, then injects action control through fine-tuning on structured annotations, maintaining both audio-visual alignment and the model's text-following capabilities. Extensive experiments demonstrate that ActAvatar significantly outperforms state-of-the-art methods in both action control and visual quality.
title ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.19546