Revisiting put-that-there, context aware window interactions via LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915595551768576 |
|---|---|
| author | Bovo, Riccardo Giunchi, Daniele Cascarano, Pasquale Gonzalez, Eric J. Gonzalez-Franco, Mar |
| author_facet | Bovo, Riccardo Giunchi, Daniele Cascarano, Pasquale Gonzalez, Eric J. Gonzalez-Franco, Mar |
| contents | We revisit Bolt's classic "Put-That-There" concept for modern head-mounted displays by pairing Large Language Models (LLMs) with XR sensor and tech stack. The agent fuses (i) a semantically segmented 3-D environment, (ii) live application metadata, and (iii) users' verbal, pointing, and head-gaze cues to issue JSON window-placement actions. As a result, users can manage a panoramic workspace through: (1) explicit commands ("Place Google Maps on the coffee table"), (2) deictic speech plus gestures ("Put that there"), or (3) high-level goals ("I need to send a message"). Unlike traditional explicit interfaces, our system supports one-to-many action mappings and goal-centric reasoning, allowing the LLM to dynamically infer relevant applications and layout decisions, including interrelationships across tools. This enables seamless, intent-driven interaction without manual window juggling in immersive XR environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_02378 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Revisiting put-that-there, context aware window interactions via LLMs Bovo, Riccardo Giunchi, Daniele Cascarano, Pasquale Gonzalez, Eric J. Gonzalez-Franco, Mar Human-Computer Interaction We revisit Bolt's classic "Put-That-There" concept for modern head-mounted displays by pairing Large Language Models (LLMs) with XR sensor and tech stack. The agent fuses (i) a semantically segmented 3-D environment, (ii) live application metadata, and (iii) users' verbal, pointing, and head-gaze cues to issue JSON window-placement actions. As a result, users can manage a panoramic workspace through: (1) explicit commands ("Place Google Maps on the coffee table"), (2) deictic speech plus gestures ("Put that there"), or (3) high-level goals ("I need to send a message"). Unlike traditional explicit interfaces, our system supports one-to-many action mappings and goal-centric reasoning, allowing the LLM to dynamically infer relevant applications and layout decisions, including interrelationships across tools. This enables seamless, intent-driven interaction without manual window juggling in immersive XR environments. |
| title | Revisiting put-that-there, context aware window interactions via LLMs |
| topic | Human-Computer Interaction |
| url | https://arxiv.org/abs/2511.02378 |