Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Jiayi, Wang, Xu, Tang, Shengeng, Wang, Yaxiong, Cheng, Lechao, Guo, Dan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909724080865280
author He, Jiayi
Wang, Xu
Tang, Shengeng
Wang, Yaxiong
Cheng, Lechao
Guo, Dan
author_facet He, Jiayi
Wang, Xu
Tang, Shengeng
Wang, Yaxiong
Cheng, Lechao
Guo, Dan
contents Sign language video generation requires producing natural signing motions with realistic appearances under precise semantic control, yet faces two critical challenges: excessive signer-specific data requirements and poor generalization. We propose a new paradigm for sign language video generation that decouples motion semantics from signer identity through a two-phase synthesis framework. First, we construct a signer-independent multimodal motion lexicon, where each gloss is stored as identity-agnostic pose, gesture, and 3D mesh sequences, requiring only one recording per sign. This compact representation enables our second key innovation: a discrete-to-continuous motion synthesis stage that transforms retrieved gloss sequences into temporally coherent motion trajectories, followed by identity-aware neural rendering to produce photorealistic videos of arbitrary signers. Unlike prior work constrained by signer-specific datasets, our method treats motion as a first-class citizen: the learned latent pose dynamics serve as a portable "choreography layer" that can be visually realized through different human appearances. Extensive experiments demonstrate that disentangling motion from identity is not just viable but advantageous - enabling both high-quality synthesis and unprecedented flexibility in signer personalization.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
He, Jiayi
Wang, Xu
Tang, Shengeng
Wang, Yaxiong
Cheng, Lechao
Guo, Dan
Computer Vision and Pattern Recognition
Sign language video generation requires producing natural signing motions with realistic appearances under precise semantic control, yet faces two critical challenges: excessive signer-specific data requirements and poor generalization. We propose a new paradigm for sign language video generation that decouples motion semantics from signer identity through a two-phase synthesis framework. First, we construct a signer-independent multimodal motion lexicon, where each gloss is stored as identity-agnostic pose, gesture, and 3D mesh sequences, requiring only one recording per sign. This compact representation enables our second key innovation: a discrete-to-continuous motion synthesis stage that transforms retrieved gloss sequences into temporally coherent motion trajectories, followed by identity-aware neural rendering to produce photorealistic videos of arbitrary signers. Unlike prior work constrained by signer-specific datasets, our method treats motion as a first-class citizen: the learned latent pose dynamics serve as a portable "choreography layer" that can be visually realized through different human appearances. Extensive experiments demonstrate that disentangling motion from identity is not just viable but advantageous - enabling both high-quality synthesis and unprecedented flexibility in signer personalization.
title Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.04049