Revisiting put-that-there, context aware window interactions via LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bovo, Riccardo, Giunchi, Daniele, Cascarano, Pasquale, Gonzalez, Eric J., Gonzalez-Franco, Mar
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915595551768576
author Bovo, Riccardo
Giunchi, Daniele
Cascarano, Pasquale
Gonzalez, Eric J.
Gonzalez-Franco, Mar
author_facet Bovo, Riccardo
Giunchi, Daniele
Cascarano, Pasquale
Gonzalez, Eric J.
Gonzalez-Franco, Mar
contents We revisit Bolt's classic "Put-That-There" concept for modern head-mounted displays by pairing Large Language Models (LLMs) with XR sensor and tech stack. The agent fuses (i) a semantically segmented 3-D environment, (ii) live application metadata, and (iii) users' verbal, pointing, and head-gaze cues to issue JSON window-placement actions. As a result, users can manage a panoramic workspace through: (1) explicit commands ("Place Google Maps on the coffee table"), (2) deictic speech plus gestures ("Put that there"), or (3) high-level goals ("I need to send a message"). Unlike traditional explicit interfaces, our system supports one-to-many action mappings and goal-centric reasoning, allowing the LLM to dynamically infer relevant applications and layout decisions, including interrelationships across tools. This enables seamless, intent-driven interaction without manual window juggling in immersive XR environments.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02378
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting put-that-there, context aware window interactions via LLMs
Bovo, Riccardo
Giunchi, Daniele
Cascarano, Pasquale
Gonzalez, Eric J.
Gonzalez-Franco, Mar
Human-Computer Interaction
We revisit Bolt's classic "Put-That-There" concept for modern head-mounted displays by pairing Large Language Models (LLMs) with XR sensor and tech stack. The agent fuses (i) a semantically segmented 3-D environment, (ii) live application metadata, and (iii) users' verbal, pointing, and head-gaze cues to issue JSON window-placement actions. As a result, users can manage a panoramic workspace through: (1) explicit commands ("Place Google Maps on the coffee table"), (2) deictic speech plus gestures ("Put that there"), or (3) high-level goals ("I need to send a message"). Unlike traditional explicit interfaces, our system supports one-to-many action mappings and goal-centric reasoning, allowing the LLM to dynamically infer relevant applications and layout decisions, including interrelationships across tools. This enables seamless, intent-driven interaction without manual window juggling in immersive XR environments.
title Revisiting put-that-there, context aware window interactions via LLMs
topic Human-Computer Interaction
url https://arxiv.org/abs/2511.02378