Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Bohong, Li, Yumeng, Zheng, Youyi, Ding, Yao-Xiang, Zhou, Kun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913961986752512
author Chen, Bohong
Li, Yumeng
Zheng, Youyi
Ding, Yao-Xiang
Zhou, Kun
author_facet Chen, Bohong
Li, Yumeng
Zheng, Youyi
Ding, Yao-Xiang
Zhou, Kun
contents The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels derived from motion examples, these approaches often compromise the rich details present in the original motion examples. We present MECo, a framework for motion-example-controlled co-speech gesture generation by leveraging large language models (LLMs). Our method capitalizes on LLMs' comprehension capabilities through fine-tuning to simultaneously interpret speech audio and motion examples, enabling the synthesis of gestures that preserve example-specific characteristics while maintaining speech congruence. Departing from conventional pseudo-labeling paradigms, we position motion examples as explicit query contexts within the prompt structure to guide gesture generation. Experimental results demonstrate state-of-the-art performance across three metrics: Fréchet Gesture Distance (FGD), motion diversity, and example-gesture similarity. Furthermore, our framework enables granular control of individual body parts and accommodates diverse input modalities including motion clips, static poses, human video sequences, and textual descriptions. Our code, pre-trained models, and videos are available at https://robinwitch.github.io/MECo-Page.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20220
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
Chen, Bohong
Li, Yumeng
Zheng, Youyi
Ding, Yao-Xiang
Zhou, Kun
Computer Vision and Pattern Recognition
The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels derived from motion examples, these approaches often compromise the rich details present in the original motion examples. We present MECo, a framework for motion-example-controlled co-speech gesture generation by leveraging large language models (LLMs). Our method capitalizes on LLMs' comprehension capabilities through fine-tuning to simultaneously interpret speech audio and motion examples, enabling the synthesis of gestures that preserve example-specific characteristics while maintaining speech congruence. Departing from conventional pseudo-labeling paradigms, we position motion examples as explicit query contexts within the prompt structure to guide gesture generation. Experimental results demonstrate state-of-the-art performance across three metrics: Fréchet Gesture Distance (FGD), motion diversity, and example-gesture similarity. Furthermore, our framework enables granular control of individual body parts and accommodates diverse input modalities including motion clips, static poses, human video sequences, and textual descriptions. Our code, pre-trained models, and videos are available at https://robinwitch.github.io/MECo-Page.
title Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.20220