LLAniMAtion: LLAMA Driven Gesture Animation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Windle, Jonathan, Matthews, Iain, Taylor, Sarah
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917664425771008
author Windle, Jonathan
Matthews, Iain
Taylor, Sarah
author_facet Windle, Jonathan
Matthews, Iain
Taylor, Sarah
contents Co-speech gesturing is an important modality in conversation, providing context and social cues. In character animation, appropriate and synchronised gestures add realism, and can make interactive agents more engaging. Historically, methods for automatically generating gestures were predominantly audio-driven, exploiting the prosodic and speech-related content that is encoded in the audio signal. In this paper we instead experiment with using LLM features for gesture generation that are extracted from text using LLAMA2. We compare against audio features, and explore combining the two modalities in both objective tests and a user study. Surprisingly, our results show that LLAMA2 features on their own perform significantly better than audio features and that including both modalities yields no significant difference to using LLAMA2 features in isolation. We demonstrate that the LLAMA2 based model can generate both beat and semantic gestures without any audio input, suggesting LLMs can provide rich encodings that are well suited for gesture generation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_08042
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLAniMAtion: LLAMA Driven Gesture Animation
Windle, Jonathan
Matthews, Iain
Taylor, Sarah
Human-Computer Interaction
Artificial Intelligence
Computer Vision and Pattern Recognition
Graphics
Machine Learning
Co-speech gesturing is an important modality in conversation, providing context and social cues. In character animation, appropriate and synchronised gestures add realism, and can make interactive agents more engaging. Historically, methods for automatically generating gestures were predominantly audio-driven, exploiting the prosodic and speech-related content that is encoded in the audio signal. In this paper we instead experiment with using LLM features for gesture generation that are extracted from text using LLAMA2. We compare against audio features, and explore combining the two modalities in both objective tests and a user study. Surprisingly, our results show that LLAMA2 features on their own perform significantly better than audio features and that including both modalities yields no significant difference to using LLAMA2 features in isolation. We demonstrate that the LLAMA2 based model can generate both beat and semantic gestures without any audio input, suggesting LLMs can provide rich encodings that are well suited for gesture generation.
title LLAniMAtion: LLAMA Driven Gesture Animation
topic Human-Computer Interaction
Artificial Intelligence
Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2405.08042