A Unified Editing Method for Co-Speech Gesture Generation via Diffusion Inversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Zeyu, Gao, Nan, Zeng, Zhi, Zhang, Guixuan, Liu, Jie, Zhang, Shuwu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910397213179904
author Zhao, Zeyu
Gao, Nan
Zeng, Zhi
Zhang, Guixuan
Liu, Jie
Zhang, Shuwu
author_facet Zhao, Zeyu
Gao, Nan
Zeng, Zhi
Zhang, Guixuan
Liu, Jie
Zhang, Shuwu
contents Diffusion models have shown great success in generating high-quality co-speech gestures for interactive humanoid robots or digital avatars from noisy input with the speech audio or text as conditions. However, they rarely focus on providing rich editing capabilities for content creators other than high-level specialized measures like style conditioning. To resolve this, we propose a unified framework utilizing diffusion inversion that enables multi-level editing capabilities for co-speech gesture generation without re-training. The method takes advantage of two key capabilities of invertible diffusion models. The first is that through inversion, we can reconstruct the intermediate noise from gestures and regenerate new gestures from the noise. This can be used to obtain gestures with high-level similarities to the original gestures for different speech conditions. The second is that this reconstruction reduces activation caching requirements during gradient calculation, making the direct optimization on input noises possible on current hardware with limited memory. With different loss functions designed for, e.g., joint rotation or velocity, we can control various low-level details by automatically tweaking the input noises through optimization. Extensive experiments on multiple use cases show that this framework succeeds in unifying high-level and low-level co-speech gesture editing.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02411
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Unified Editing Method for Co-Speech Gesture Generation via Diffusion Inversion
Zhao, Zeyu
Gao, Nan
Zeng, Zhi
Zhang, Guixuan
Liu, Jie
Zhang, Shuwu
Human-Computer Interaction
Diffusion models have shown great success in generating high-quality co-speech gestures for interactive humanoid robots or digital avatars from noisy input with the speech audio or text as conditions. However, they rarely focus on providing rich editing capabilities for content creators other than high-level specialized measures like style conditioning. To resolve this, we propose a unified framework utilizing diffusion inversion that enables multi-level editing capabilities for co-speech gesture generation without re-training. The method takes advantage of two key capabilities of invertible diffusion models. The first is that through inversion, we can reconstruct the intermediate noise from gestures and regenerate new gestures from the noise. This can be used to obtain gestures with high-level similarities to the original gestures for different speech conditions. The second is that this reconstruction reduces activation caching requirements during gradient calculation, making the direct optimization on input noises possible on current hardware with limited memory. With different loss functions designed for, e.g., joint rotation or velocity, we can control various low-level details by automatically tweaking the input noises through optimization. Extensive experiments on multiple use cases show that this framework succeeds in unifying high-level and low-level co-speech gesture editing.
title A Unified Editing Method for Co-Speech Gesture Generation via Diffusion Inversion
topic Human-Computer Interaction
url https://arxiv.org/abs/2404.02411