Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guichoux, Téo, Lemerle, Théodor, Mehta, Shivam, Beskow, Jonas, Henter, Gustav Eje, Soulier, Laure, Pelachaud, Catherine, Obin, Nicolas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915961598115840
author Guichoux, Téo
Lemerle, Théodor
Mehta, Shivam
Beskow, Jonas
Henter, Gustav Eje
Soulier, Laure
Pelachaud, Catherine
Obin, Nicolas
author_facet Guichoux, Téo
Lemerle, Théodor
Mehta, Shivam
Beskow, Jonas
Henter, Gustav Eje
Soulier, Laure
Pelachaud, Catherine
Obin, Nicolas
contents Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-speech gestures from text using interleaved token sequences in a discrete autoregressive backbone, with modality-specific decoders. Gelina supports multi-speaker and multi-style cloning and enables gesture-only synthesis from speech inputs. Subjective and objective evaluations demonstrate competitive speech quality and improved gesture generation over unimodal baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12834
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
Guichoux, Téo
Lemerle, Théodor
Mehta, Shivam
Beskow, Jonas
Henter, Gustav Eje
Soulier, Laure
Pelachaud, Catherine
Obin, Nicolas
Sound
Artificial Intelligence
Audio and Speech Processing
68T07
Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-speech gestures from text using interleaved token sequences in a discrete autoregressive backbone, with modality-specific decoders. Gelina supports multi-speaker and multi-style cloning and enables gesture-only synthesis from speech inputs. Subjective and objective evaluations demonstrate competitive speech quality and improved gesture generation over unimodal baselines.
title Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
topic Sound
Artificial Intelligence
Audio and Speech Processing
68T07
url https://arxiv.org/abs/2510.12834