Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mughal, M. Hamza, Dabral, Rishabh, Scholman, Merel C. J., Demberg, Vera, Theobalt, Christian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916673768914944
author Mughal, M. Hamza
Dabral, Rishabh
Scholman, Merel C. J.
Demberg, Vera
Theobalt, Christian
author_facet Mughal, M. Hamza
Dabral, Rishabh
Scholman, Merel C. J.
Demberg, Vera
Theobalt, Christian
contents Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate rhythmic beat gestures, but struggle to produce semantically meaningful gestures. Therefore, we present RAG-Gesture, a diffusion-based gesture generation approach that leverages Retrieval Augmented Generation (RAG) to produce natural-looking and semantically rich gestures. Our neuro-explicit gesture generation approach is designed to produce semantic gestures grounded in interpretable linguistic knowledge. We achieve this by using explicit domain knowledge to retrieve exemplar motions from a database of co-speech gestures. Once retrieved, we then inject these semantic exemplar gestures into our diffusion-based gesture generation pipeline using DDIM inversion and retrieval guidance at the inference time without any need of training. Further, we propose a control paradigm for guidance, that allows the users to modulate the amount of influence each retrieval insertion has over the generated sequence. Our comparative evaluations demonstrate the validity of our approach against recent gesture generation approaches. The reader is urged to explore the results on our project page.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06786
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis
Mughal, M. Hamza
Dabral, Rishabh
Scholman, Merel C. J.
Demberg, Vera
Theobalt, Christian
Computer Vision and Pattern Recognition
Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate rhythmic beat gestures, but struggle to produce semantically meaningful gestures. Therefore, we present RAG-Gesture, a diffusion-based gesture generation approach that leverages Retrieval Augmented Generation (RAG) to produce natural-looking and semantically rich gestures. Our neuro-explicit gesture generation approach is designed to produce semantic gestures grounded in interpretable linguistic knowledge. We achieve this by using explicit domain knowledge to retrieve exemplar motions from a database of co-speech gestures. Once retrieved, we then inject these semantic exemplar gestures into our diffusion-based gesture generation pipeline using DDIM inversion and retrieval guidance at the inference time without any need of training. Further, we propose a control paradigm for guidance, that allows the users to modulate the amount of influence each retrieval insertion has over the generated sequence. Our comparative evaluations demonstrate the validity of our approach against recent gesture generation approaches. The reader is urged to explore the results on our project page.
title Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.06786