Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AI

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Chengyuan, Kumaran, Radha, Stier, Noah, Yu, Kangyou, Höllerer, Tobias
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929530019512320
author Xu, Chengyuan
Kumaran, Radha
Stier, Noah
Yu, Kangyou
Höllerer, Tobias
author_facet Xu, Chengyuan
Kumaran, Radha
Stier, Noah
Yu, Kangyou
Höllerer, Tobias
contents Seamless integration of virtual and physical worlds in augmented reality benefits from the system semantically "understanding" the physical environment. AR research has long focused on the potential of context awareness, demonstrating novel capabilities that leverage the semantics in the 3D environment for various object-level interactions. Meanwhile, the computer vision community has made leaps in neural vision-language understanding to enhance environment perception for autonomous tasks. In this work, we introduce a multimodal 3D object representation that unifies both semantic and linguistic knowledge with the geometric representation, enabling user-guided machine learning involving physical objects. We first present a fast multimodal 3D reconstruction pipeline that brings linguistic understanding to AR by fusing CLIP vision-language features into the environment and object models. We then propose "in-situ" machine learning, which, in conjunction with the multimodal representation, enables new tools and interfaces for users to interact with physical spaces and objects in a spatially and linguistically meaningful manner. We demonstrate the usefulness of the proposed system through two real-world AR applications on Magic Leap 2: a) spatial search in physical environments with natural language and b) an intelligent inventory system that tracks object changes over time. We also make our full implementation and demo data available at (https://github.com/cy-xu/spatially_aware_AI) to encourage further exploration and research in spatially aware AI.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04652
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AI
Xu, Chengyuan
Kumaran, Radha
Stier, Noah
Yu, Kangyou
Höllerer, Tobias
Human-Computer Interaction
Artificial Intelligence
Computer Vision and Pattern Recognition
I.4.8; H.5.2
Seamless integration of virtual and physical worlds in augmented reality benefits from the system semantically "understanding" the physical environment. AR research has long focused on the potential of context awareness, demonstrating novel capabilities that leverage the semantics in the 3D environment for various object-level interactions. Meanwhile, the computer vision community has made leaps in neural vision-language understanding to enhance environment perception for autonomous tasks. In this work, we introduce a multimodal 3D object representation that unifies both semantic and linguistic knowledge with the geometric representation, enabling user-guided machine learning involving physical objects. We first present a fast multimodal 3D reconstruction pipeline that brings linguistic understanding to AR by fusing CLIP vision-language features into the environment and object models. We then propose "in-situ" machine learning, which, in conjunction with the multimodal representation, enables new tools and interfaces for users to interact with physical spaces and objects in a spatially and linguistically meaningful manner. We demonstrate the usefulness of the proposed system through two real-world AR applications on Magic Leap 2: a) spatial search in physical environments with natural language and b) an intelligent inventory system that tracks object changes over time. We also make our full implementation and demo data available at (https://github.com/cy-xu/spatially_aware_AI) to encourage further exploration and research in spatially aware AI.
title Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AI
topic Human-Computer Interaction
Artificial Intelligence
Computer Vision and Pattern Recognition
I.4.8; H.5.2
url https://arxiv.org/abs/2410.04652