Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915961598115840 |
|---|---|
| author | Guichoux, Téo Lemerle, Théodor Mehta, Shivam Beskow, Jonas Henter, Gustav Eje Soulier, Laure Pelachaud, Catherine Obin, Nicolas |
| author_facet | Guichoux, Téo Lemerle, Théodor Mehta, Shivam Beskow, Jonas Henter, Gustav Eje Soulier, Laure Pelachaud, Catherine Obin, Nicolas |
| contents | Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-speech gestures from text using interleaved token sequences in a discrete autoregressive backbone, with modality-specific decoders. Gelina supports multi-speaker and multi-style cloning and enables gesture-only synthesis from speech inputs. Subjective and objective evaluations demonstrate competitive speech quality and improved gesture generation over unimodal baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_12834 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction Guichoux, Téo Lemerle, Théodor Mehta, Shivam Beskow, Jonas Henter, Gustav Eje Soulier, Laure Pelachaud, Catherine Obin, Nicolas Sound Artificial Intelligence Audio and Speech Processing 68T07 Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-speech gestures from text using interleaved token sequences in a discrete autoregressive backbone, with modality-specific decoders. Gelina supports multi-speaker and multi-style cloning and enables gesture-only synthesis from speech inputs. Subjective and objective evaluations demonstrate competitive speech quality and improved gesture generation over unimodal baselines. |
| title | Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction |
| topic | Sound Artificial Intelligence Audio and Speech Processing 68T07 |
| url | https://arxiv.org/abs/2510.12834 |