SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929761058553856 |
|---|---|
| author | Wang, Wenjia Pan, Liang Dou, Zhiyang Mei, Jidong Liao, Zhouyingcheng Lou, Yuke Wu, Yifan Yang, Lei Wang, Jingbo Komura, Taku |
| author_facet | Wang, Wenjia Pan, Liang Dou, Zhiyang Mei, Jidong Liao, Zhouyingcheng Lou, Yuke Wu, Yifan Yang, Lei Wang, Jingbo Komura, Taku |
| contents | Simulating stylized human-scene interactions (HSI) in physical environments is a challenging yet fascinating task. Prior works emphasize long-term execution but fall short in achieving both diverse style and physical plausibility. To tackle this challenge, we introduce a novel hierarchical framework named SIMS that seamlessly bridges highlevel script-driven intent with a low-level control policy, enabling more expressive and diverse human-scene interactions. Specifically, we employ Large Language Models with Retrieval-Augmented Generation (RAG) to generate coherent and diverse long-form scripts, providing a rich foundation for motion planning. A versatile multicondition physics-based control policy is also developed, which leverages text embeddings from the generated scripts to encode stylistic cues, simultaneously perceiving environmental geometries and accomplishing task goals. By integrating the retrieval-augmented script generation with the multi-condition controller, our approach provides a unified solution for generating stylized HSI motions. We further introduce a comprehensive planning dataset produced by RAG and a stylized motion dataset featuring diverse locomotions and interactions. Extensive experiments demonstrate SIMS's effectiveness in executing various tasks and generalizing across different scenarios, significantly outperforming previous methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_19921 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation Wang, Wenjia Pan, Liang Dou, Zhiyang Mei, Jidong Liao, Zhouyingcheng Lou, Yuke Wu, Yifan Yang, Lei Wang, Jingbo Komura, Taku Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Graphics Simulating stylized human-scene interactions (HSI) in physical environments is a challenging yet fascinating task. Prior works emphasize long-term execution but fall short in achieving both diverse style and physical plausibility. To tackle this challenge, we introduce a novel hierarchical framework named SIMS that seamlessly bridges highlevel script-driven intent with a low-level control policy, enabling more expressive and diverse human-scene interactions. Specifically, we employ Large Language Models with Retrieval-Augmented Generation (RAG) to generate coherent and diverse long-form scripts, providing a rich foundation for motion planning. A versatile multicondition physics-based control policy is also developed, which leverages text embeddings from the generated scripts to encode stylistic cues, simultaneously perceiving environmental geometries and accomplishing task goals. By integrating the retrieval-augmented script generation with the multi-condition controller, our approach provides a unified solution for generating stylized HSI motions. We further introduce a comprehensive planning dataset produced by RAG and a stylized motion dataset featuring diverse locomotions and interactions. Extensive experiments demonstrate SIMS's effectiveness in executing various tasks and generalizing across different scenarios, significantly outperforming previous methods. |
| title | SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Graphics |
| url | https://arxiv.org/abs/2411.19921 |