ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Khan, Dawar, Kouyoumdjian, Alexandre, Liu, Xinyu, Mena, Omar, Engel, Dominik, Viola, Ivan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911570032852992
author Khan, Dawar
Kouyoumdjian, Alexandre
Liu, Xinyu
Mena, Omar
Engel, Dominik
Viola, Ivan
author_facet Khan, Dawar
Kouyoumdjian, Alexandre
Liu, Xinyu
Mena, Omar
Engel, Dominik
Viola, Ivan
contents We present ClickAIXR, a novel on-device framework for multimodal vision-language interaction with objects in extended reality (XR). Unlike prior systems that rely on cloud-based AI (e.g., ChatGPT) or gaze-based selection (e.g., GazePointAR), ClickAIXR integrates an on-device vision-language model (VLM) with a controller-based object selection paradigm, enabling users to precisely click on real-world objects in XR. Once selected, the object image is processed locally by the VLM to answer natural language questions through both text and speech. This object-centered interaction reduces ambiguity inherent in gaze- or voice-only interfaces and improves transparency by performing all inference on-device, addressing concerns around privacy and latency. We implemented ClickAIXR in the Magic Leap SDK (C API) with ONNX-based local VLM inference. We conducted a user study comparing ClickAIXR with Gemini 2.5 Flash and ChatGPT 5, evaluating usability, trust, and user satisfaction. Results show that latency is moderate and user experience is acceptable. Our findings demonstrate the potential of click-based object selection combined with on-device AI to advance trustworthy, privacy-preserving XR interactions. The source code and supplementary materials are available at: nanovis.org/ClickAIXR.html
format Preprint
id arxiv_https___arxiv_org_abs_2604_04905
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality
Khan, Dawar
Kouyoumdjian, Alexandre
Liu, Xinyu
Mena, Omar
Engel, Dominik
Viola, Ivan
Computer Vision and Pattern Recognition
Graphics
Human-Computer Interaction
We present ClickAIXR, a novel on-device framework for multimodal vision-language interaction with objects in extended reality (XR). Unlike prior systems that rely on cloud-based AI (e.g., ChatGPT) or gaze-based selection (e.g., GazePointAR), ClickAIXR integrates an on-device vision-language model (VLM) with a controller-based object selection paradigm, enabling users to precisely click on real-world objects in XR. Once selected, the object image is processed locally by the VLM to answer natural language questions through both text and speech. This object-centered interaction reduces ambiguity inherent in gaze- or voice-only interfaces and improves transparency by performing all inference on-device, addressing concerns around privacy and latency. We implemented ClickAIXR in the Magic Leap SDK (C API) with ONNX-based local VLM inference. We conducted a user study comparing ClickAIXR with Gemini 2.5 Flash and ChatGPT 5, evaluating usability, trust, and user satisfaction. Results show that latency is moderate and user experience is acceptable. Our findings demonstrate the potential of click-based object selection combined with on-device AI to advance trustworthy, privacy-preserving XR interactions. The source code and supplementary materials are available at: nanovis.org/ClickAIXR.html
title ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality
topic Computer Vision and Pattern Recognition
Graphics
Human-Computer Interaction
url https://arxiv.org/abs/2604.04905