MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Shurong, Li, Huadong, Wu, Juhao, Jing, Minhao, Li, Linze, Ji, Renhe, Liang, Jiajun, Fan, Haoqiang, Wang, Jin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909297566285824
author Yang, Shurong
Li, Huadong
Wu, Juhao
Jing, Minhao
Li, Linze
Ji, Renhe
Liang, Jiajun
Fan, Haoqiang
Wang, Jin
author_facet Yang, Shurong
Li, Huadong
Wu, Juhao
Jing, Minhao
Li, Linze
Ji, Renhe
Liang, Jiajun
Fan, Haoqiang
Wang, Jin
contents Diffusion models have demonstrated superior performance in the field of portrait animation. However, current approaches relied on either visual or audio modality to control character movements, failing to exploit the potential of mixed-modal control. This challenge arises from the difficulty in balancing the weak control strength of audio modality and the strong control strength of visual modality. To address this issue, we introduce MegActor-$Σ$: a mixed-modal conditional diffusion transformer (DiT), which can flexibly inject audio and visual modality control signals into portrait animation. Specifically, we make substantial advancements over its predecessor, MegActor, by leveraging the promising model structure of DiT and integrating audio and visual conditions through advanced modules within the DiT framework. To further achieve flexible combinations of mixed-modal control signals, we propose a ``Modality Decoupling Control" training strategy to balance the control strength between visual and audio modalities, along with the ``Amplitude Adjustment" inference strategy to freely regulate the motion amplitude of each modality. Finally, to facilitate extensive studies in this field, we design several dataset evaluation metrics to filter out public datasets and solely use this filtered dataset to train MegActor-$Σ$. Extensive experiments demonstrate the superiority of our approach in generating vivid portrait animations, outperforming previous methods trained on private dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14975
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer
Yang, Shurong
Li, Huadong
Wu, Juhao
Jing, Minhao
Li, Linze
Ji, Renhe
Liang, Jiajun
Fan, Haoqiang
Wang, Jin
Computer Vision and Pattern Recognition
Diffusion models have demonstrated superior performance in the field of portrait animation. However, current approaches relied on either visual or audio modality to control character movements, failing to exploit the potential of mixed-modal control. This challenge arises from the difficulty in balancing the weak control strength of audio modality and the strong control strength of visual modality. To address this issue, we introduce MegActor-$Σ$: a mixed-modal conditional diffusion transformer (DiT), which can flexibly inject audio and visual modality control signals into portrait animation. Specifically, we make substantial advancements over its predecessor, MegActor, by leveraging the promising model structure of DiT and integrating audio and visual conditions through advanced modules within the DiT framework. To further achieve flexible combinations of mixed-modal control signals, we propose a ``Modality Decoupling Control" training strategy to balance the control strength between visual and audio modalities, along with the ``Amplitude Adjustment" inference strategy to freely regulate the motion amplitude of each modality. Finally, to facilitate extensive studies in this field, we design several dataset evaluation metrics to filter out public datasets and solely use this filtered dataset to train MegActor-$Σ$. Extensive experiments demonstrate the superiority of our approach in generating vivid portrait animations, outperforming previous methods trained on private dataset.
title MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.14975