MOSPA: Human Motion Generation Driven by Spatial Audio

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Shuyang, Dou, Zhiyang, Shi, Mingyi, Pan, Liang, Ho, Leo, Wang, Jingbo, Liu, Yuan, Lin, Cheng, Ma, Yuexin, Wang, Wenping, Komura, Taku
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909881444859904
author Xu, Shuyang
Dou, Zhiyang
Shi, Mingyi
Pan, Liang
Ho, Leo
Wang, Jingbo
Liu, Yuan
Lin, Cheng
Ma, Yuexin
Wang, Wenping
Komura, Taku
author_facet Xu, Shuyang
Dou, Zhiyang
Shi, Mingyi
Pan, Liang
Ho, Leo
Wang, Jingbo
Liu, Yuan
Lin, Cheng
Ma, Yuexin
Wang, Wenping
Komura, Taku
contents Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have primarily focused on mapping modalities like speech, audio, and music to generate human motion. As of yet, these models typically overlook the impact of spatial features encoded in spatial audio signals on human motion. To bridge this gap and enable high-quality modeling of human movements in response to spatial audio, we introduce the first comprehensive Spatial Audio-Driven Human Motion (SAM) dataset, which contains diverse and high-quality spatial audio and motion data. For benchmarking, we develop a simple yet effective diffusion-based generative framework for human MOtion generation driven by SPatial Audio, termed MOSPA, which faithfully captures the relationship between body motion and spatial audio through an effective fusion mechanism. Once trained, MOSPA can generate diverse, realistic human motions conditioned on varying spatial audio inputs. We perform a thorough investigation of the proposed dataset and conduct extensive experiments for benchmarking, where our method achieves state-of-the-art performance on this task. Our code and model are publicly available at https://github.com/xsy27/Mospa-Acoustic-driven-Motion-Generation
format Preprint
id arxiv_https___arxiv_org_abs_2507_11949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MOSPA: Human Motion Generation Driven by Spatial Audio
Xu, Shuyang
Dou, Zhiyang
Shi, Mingyi
Pan, Liang
Ho, Leo
Wang, Jingbo
Liu, Yuan
Lin, Cheng
Ma, Yuexin
Wang, Wenping
Komura, Taku
Graphics
Computer Vision and Pattern Recognition
Robotics
Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have primarily focused on mapping modalities like speech, audio, and music to generate human motion. As of yet, these models typically overlook the impact of spatial features encoded in spatial audio signals on human motion. To bridge this gap and enable high-quality modeling of human movements in response to spatial audio, we introduce the first comprehensive Spatial Audio-Driven Human Motion (SAM) dataset, which contains diverse and high-quality spatial audio and motion data. For benchmarking, we develop a simple yet effective diffusion-based generative framework for human MOtion generation driven by SPatial Audio, termed MOSPA, which faithfully captures the relationship between body motion and spatial audio through an effective fusion mechanism. Once trained, MOSPA can generate diverse, realistic human motions conditioned on varying spatial audio inputs. We perform a thorough investigation of the proposed dataset and conduct extensive experiments for benchmarking, where our method achieves state-of-the-art performance on this task. Our code and model are publicly available at https://github.com/xsy27/Mospa-Acoustic-driven-Motion-Generation
title MOSPA: Human Motion Generation Driven by Spatial Audio
topic Graphics
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2507.11949