Crafting Dynamic Virtual Activities with Advanced Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Changyang, Yan, Qingan, Kim, Minyoung, Li, Zhan, Xu, Yi, Yu, Lap-Fai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909898731683840
author Li, Changyang
Yan, Qingan
Kim, Minyoung
Li, Zhan
Xu, Yi
Yu, Lap-Fai
author_facet Li, Changyang
Yan, Qingan
Kim, Minyoung
Li, Zhan
Xu, Yi
Yu, Lap-Fai
contents In this paper, we investigate the use of multimodal large language models (MLLMs) for generating virtual activities, leveraging the integration of vision-language modalities to enable the interpretation of virtual environments. Our approach recognizes and abstracts key scene elements including scene layouts, semantic contexts, and object identities with MLLMs' multimodal reasoning capabilities. By correlating these abstractions with massive knowledge about human activities, MLLMs are capable of generating adaptive and contextually relevant virtual activities. We propose a structured framework to articulate abstract activity descriptions, emphasizing detailed multi-character interactions within virtual spaces. Utilizing the derived high-level contexts, our approach accurately positions virtual characters and ensures that their interactions and behaviors are realistically and contextually appropriate through strategic optimization. Experiment results demonstrate the effectiveness of our approach, providing a novel direction for enhancing the realism and context-awareness in simulated virtual environments.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17582
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Crafting Dynamic Virtual Activities with Advanced Multimodal Models
Li, Changyang
Yan, Qingan
Kim, Minyoung
Li, Zhan
Xu, Yi
Yu, Lap-Fai
Human-Computer Interaction
Graphics
Multimedia
In this paper, we investigate the use of multimodal large language models (MLLMs) for generating virtual activities, leveraging the integration of vision-language modalities to enable the interpretation of virtual environments. Our approach recognizes and abstracts key scene elements including scene layouts, semantic contexts, and object identities with MLLMs' multimodal reasoning capabilities. By correlating these abstractions with massive knowledge about human activities, MLLMs are capable of generating adaptive and contextually relevant virtual activities. We propose a structured framework to articulate abstract activity descriptions, emphasizing detailed multi-character interactions within virtual spaces. Utilizing the derived high-level contexts, our approach accurately positions virtual characters and ensures that their interactions and behaviors are realistically and contextually appropriate through strategic optimization. Experiment results demonstrate the effectiveness of our approach, providing a novel direction for enhancing the realism and context-awareness in simulated virtual environments.
title Crafting Dynamic Virtual Activities with Advanced Multimodal Models
topic Human-Computer Interaction
Graphics
Multimedia
url https://arxiv.org/abs/2406.17582