Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Pinxin, Zhang, Pengfei, Kim, Hyeongwoo, Garrido, Pablo, Shapiro, Ari, Olszewski, Kyle
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918145232470016
author Liu, Pinxin
Zhang, Pengfei
Kim, Hyeongwoo
Garrido, Pablo
Shapiro, Ari
Olszewski, Kyle
author_facet Liu, Pinxin
Zhang, Pengfei
Kim, Hyeongwoo
Garrido, Pablo
Shapiro, Ari
Olszewski, Kyle
contents Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accurately identifying the rhythmic or semantic triggers from audio for generating contextualized gesture patterns and achieving pixel-level realism. To address these challenges, we introduce Contextual Gesture, a framework that improves co-speech gesture video generation through three innovative components: (1) a chronological speech-gesture alignment that temporally connects two modalities, (2) a contextualized gesture tokenization that incorporate speech context into motion pattern representation through distillation, and (3) a structure-aware refinement module that employs edge connection to link gesture keypoints to improve video generation. Our extensive experiments demonstrate that Contextual Gesture not only produces realistic and speech-aligned gesture videos but also supports long-sequence generation and video gesture editing applications, shown in Fig.1.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07239
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
Liu, Pinxin
Zhang, Pengfei
Kim, Hyeongwoo
Garrido, Pablo
Shapiro, Ari
Olszewski, Kyle
Computer Vision and Pattern Recognition
Artificial Intelligence
Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accurately identifying the rhythmic or semantic triggers from audio for generating contextualized gesture patterns and achieving pixel-level realism. To address these challenges, we introduce Contextual Gesture, a framework that improves co-speech gesture video generation through three innovative components: (1) a chronological speech-gesture alignment that temporally connects two modalities, (2) a contextualized gesture tokenization that incorporate speech context into motion pattern representation through distillation, and (3) a structure-aware refinement module that employs edge connection to link gesture keypoints to improve video generation. Our extensive experiments demonstrate that Contextual Gesture not only produces realistic and speech-aligned gesture videos but also supports long-sequence generation and video gesture editing applications, shown in Fig.1.
title Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2502.07239