SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Jianping, Xiao, Weiye, Lin, Zhengyu, Zhang, Huaizhong, Ren, Tianxiang, Gao, Yang, Lin, Zhiqian, Cai, Zhongang, Yang, Lei, Liu, Ziwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909410236825600
author Jiang, Jianping
Xiao, Weiye
Lin, Zhengyu
Zhang, Huaizhong
Ren, Tianxiang
Gao, Yang
Lin, Zhiqian
Cai, Zhongang
Yang, Lei
Liu, Ziwei
author_facet Jiang, Jianping
Xiao, Weiye
Lin, Zhengyu
Zhang, Huaizhong
Ren, Tianxiang
Gao, Yang
Lin, Zhiqian
Cai, Zhongang
Yang, Lei
Liu, Ziwei
contents Human beings are social animals. How to equip 3D autonomous characters with similar social intelligence that can perceive, understand and interact with humans remains an open yet foundamental problem. In this paper, we introduce SOLAMI, the first end-to-end Social vision-Language-Action (VLA) Modeling framework for Immersive interaction with 3D autonomous characters. Specifically, SOLAMI builds 3D autonomous characters from three aspects: (1) Social VLA Architecture: We propose a unified social VLA framework to generate multimodal response (speech and motion) based on the user's multimodal input to drive the character for social interaction. (2) Interactive Multimodal Data: We present SynMSI, a synthetic multimodal social interaction dataset generated by an automatic pipeline using only existing motion datasets to address the issue of data scarcity. (3) Immersive VR Interface: We develop a VR interface that enables users to immersively interact with these characters driven by various architectures. Extensive quantitative experiments and user studies demonstrate that our framework leads to more precise and natural character responses (in both speech and motion) that align with user expectations with lower latency.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00174
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters
Jiang, Jianping
Xiao, Weiye
Lin, Zhengyu
Zhang, Huaizhong
Ren, Tianxiang
Gao, Yang
Lin, Zhiqian
Cai, Zhongang
Yang, Lei
Liu, Ziwei
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Human beings are social animals. How to equip 3D autonomous characters with similar social intelligence that can perceive, understand and interact with humans remains an open yet foundamental problem. In this paper, we introduce SOLAMI, the first end-to-end Social vision-Language-Action (VLA) Modeling framework for Immersive interaction with 3D autonomous characters. Specifically, SOLAMI builds 3D autonomous characters from three aspects: (1) Social VLA Architecture: We propose a unified social VLA framework to generate multimodal response (speech and motion) based on the user's multimodal input to drive the character for social interaction. (2) Interactive Multimodal Data: We present SynMSI, a synthetic multimodal social interaction dataset generated by an automatic pipeline using only existing motion datasets to address the issue of data scarcity. (3) Immersive VR Interface: We develop a VR interface that enables users to immersively interact with these characters driven by various architectures. Extensive quantitative experiments and user studies demonstrate that our framework leads to more precise and natural character responses (in both speech and motion) that align with user expectations with lower latency.
title SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.00174