Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Jingyu, Gong, Kunkun, Tong, Zhuoran, Chen, Chuanhan, Yuan, Mingang, Chen, Zhizhong, Zhang, Xin, Tan, Yuan, Xie
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917072918806528
author Jingyu, Gong
Kunkun, Tong
Zhuoran, Chen
Chuanhan, Yuan
Mingang, Chen
Zhizhong, Zhang
Xin, Tan
Yuan, Xie
author_facet Jingyu, Gong
Kunkun, Tong
Zhuoran, Chen
Chuanhan, Yuan
Mingang, Chen
Zhizhong, Zhang
Xin, Tan
Yuan, Xie
contents Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take an unified Scene Semantic Occupancy (SSO) for scene representation, termed SSOMotion. We design a bi-directional tri-plane decomposition to derive a compact version of the SSO, and scene semantics are mapped to an unified feature space via CLIP encoding and shared linear dimensionality reduction. Such strategy can derive the fine-grained scene semantic structures while significantly reduce redundant computations. We further take these scene hints and movement direction derived from instructions for motion control via frame-wise scene query. Extensive experiments and ablation studies conducted on cluttered scenes using ShapeNet furniture, as well as scanned scenes from PROX and Replica datasets, demonstrate its cutting-edge performance while validating its effectiveness and generalization ability. Code will be publicly available at https://github.com/jingyugong/SSOMotion.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07819
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
Jingyu, Gong
Kunkun, Tong
Zhuoran, Chen
Chuanhan, Yuan
Mingang, Chen
Zhizhong, Zhang
Xin, Tan
Yuan, Xie
Computer Vision and Pattern Recognition
Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take an unified Scene Semantic Occupancy (SSO) for scene representation, termed SSOMotion. We design a bi-directional tri-plane decomposition to derive a compact version of the SSO, and scene semantics are mapped to an unified feature space via CLIP encoding and shared linear dimensionality reduction. Such strategy can derive the fine-grained scene semantic structures while significantly reduce redundant computations. We further take these scene hints and movement direction derived from instructions for motion control via frame-wise scene query. Extensive experiments and ablation studies conducted on cluttered scenes using ShapeNet furniture, as well as scanned scenes from PROX and Replica datasets, demonstrate its cutting-edge performance while validating its effectiveness and generalization ability. Code will be publicly available at https://github.com/jingyugong/SSOMotion.
title Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.07819