Words into World: A Task-Adaptive Agent for Language-Guided Spatial Retrieval in AR
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Lixing, Höllerer, Tobias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comparing Zealous and Restrained AI Recommendations in a Real-World Human-AI Collaboration Task
by: Xu, Chengyuan, et al.
Published: (2024)
by: Xu, Chengyuan, et al.
Published: (2024)
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
by: Duan, Lin, et al.
Published: (2025)
by: Duan, Lin, et al.
Published: (2025)
Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AI
by: Xu, Chengyuan, et al.
Published: (2024)
by: Xu, Chengyuan, et al.
Published: (2024)
ScreenAgent: A Vision Language Model-driven Computer Control Agent
by: Niu, Runliang, et al.
Published: (2024)
by: Niu, Runliang, et al.
Published: (2024)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
Encode-Store-Retrieve: Augmenting Human Memory through Language-Encoded Egocentric Perception
by: Shen, Junxiao, et al.
Published: (2023)
by: Shen, Junxiao, et al.
Published: (2023)
Do Vision Language Models Understand Human Engagement in Games?
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
by: Cheng, Kanzhi, et al.
Published: (2026)
by: Cheng, Kanzhi, et al.
Published: (2026)
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
by: Yang, Saelyne, et al.
Published: (2026)
by: Yang, Saelyne, et al.
Published: (2026)
Yume: An Interactive World Generation Model
by: Mao, Xiaofeng, et al.
Published: (2025)
by: Mao, Xiaofeng, et al.
Published: (2025)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
by: Wu, Zongru, et al.
Published: (2025)
by: Wu, Zongru, et al.
Published: (2025)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
by: Sun, Qiushi, et al.
Published: (2024)
by: Sun, Qiushi, et al.
Published: (2024)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
by: Natalie, Rosiana, et al.
Published: (2025)
by: Natalie, Rosiana, et al.
Published: (2025)
A Picture is Worth a Thousand Prompts? Efficacy of Iterative Human-Driven Prompt Refinement in Image Regeneration Tasks
by: Trinh, Khoi, et al.
Published: (2025)
by: Trinh, Khoi, et al.
Published: (2025)
AutoTour: Automatic Photo Tour Guide with Smartphones and LLMs
by: Xu, Huatao, et al.
Published: (2026)
by: Xu, Huatao, et al.
Published: (2026)
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
by: Zhao, Haoyu, et al.
Published: (2025)
by: Zhao, Haoyu, et al.
Published: (2025)
SASG-DA: Sparse-Aware Semantic-Guided Diffusion Augmentation For Myoelectric Gesture Recognition
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
by: Kondic, Jovana, et al.
Published: (2025)
by: Kondic, Jovana, et al.
Published: (2025)
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
by: Hu, Haobo, et al.
Published: (2026)
by: Hu, Haobo, et al.
Published: (2026)
Adaptive 3D UI Placement in Mixed Reality Using Deep Reinforcement Learning
by: Lu, Feiyu, et al.
Published: (2025)
by: Lu, Feiyu, et al.
Published: (2025)
Regressor-Guided Generative Image Editing Balances User Emotions to Reduce Time Spent Online
by: Gebhardt, Christoph, et al.
Published: (2025)
by: Gebhardt, Christoph, et al.
Published: (2025)
HandS3C: 3D Hand Mesh Reconstruction with State Space Spatial Channel Attention from RGB images
by: Jiao, Zixun, et al.
Published: (2024)
by: Jiao, Zixun, et al.
Published: (2024)
A Backbone for Long-Horizon Robot Task Understanding
by: Chen, Xiaoshuai, et al.
Published: (2024)
by: Chen, Xiaoshuai, et al.
Published: (2024)
T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
by: Chen, Chieh-Yun, et al.
Published: (2025)
by: Chen, Chieh-Yun, et al.
Published: (2025)
GUICourse: From General Vision Language Models to Versatile GUI Agents
by: Chen, Wentong, et al.
Published: (2024)
by: Chen, Wentong, et al.
Published: (2024)
Code2World: A GUI World Model via Renderable Code Generation
by: Zheng, Yuhao, et al.
Published: (2026)
by: Zheng, Yuhao, et al.
Published: (2026)
BdSLW401: Transformer-Based Word-Level Bangla Sign Language Recognition Using Relative Quantization Encoding (RQE)
by: Rubaiyeat, Husne Ara, et al.
Published: (2025)
by: Rubaiyeat, Husne Ara, et al.
Published: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Forest-Chat: Adapting Vision-Language Agents for Interactive Forest Change Analysis
by: Brock, James, et al.
Published: (2026)
by: Brock, James, et al.
Published: (2026)
Trust in Vision-Language Models: Insights from a Participatory User Workshop
by: Chiatti, Agnese, et al.
Published: (2025)
by: Chiatti, Agnese, et al.
Published: (2025)
Large Language Models estimate fine-grained human color-concept associations
by: Mukherjee, Kushin, et al.
Published: (2024)
by: Mukherjee, Kushin, et al.
Published: (2024)
VLM-driven Behavior Tree for Context-aware Task Planning
by: Wake, Naoki, et al.
Published: (2025)
by: Wake, Naoki, et al.
Published: (2025)
OLMD: Orientation-aware Long-term Motion Decoupling for Continuous Sign Language Recognition
by: Yu, Yiheng, et al.
Published: (2025)
by: Yu, Yiheng, et al.
Published: (2025)
Constructive Apraxia: An Unexpected Limit of Instructible Vision-Language Models and Analog for Human Cognitive Disorders
by: Noever, David, et al.
Published: (2024)
by: Noever, David, et al.
Published: (2024)
Achieving Effective Virtual Reality Interactions via Acoustic Gesture Recognition based on Large Language Models
by: Zhang, Xijie, et al.
Published: (2025)
by: Zhang, Xijie, et al.
Published: (2025)
PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
UI-UG: A Unified MLLM for UI Understanding and Generation
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
by: Sun, Boyuan, et al.
Published: (2026)
by: Sun, Boyuan, et al.
Published: (2026)
ImageTalk: Designing a Multimodal AAC Text Generation System Driven by Image Recognition and Natural Language Generation
by: Yang, Boyin, et al.
Published: (2025)
by: Yang, Boyin, et al.
Published: (2025)
VerSe: Integrating Multiple Queries as Prompts for Versatile Cardiac MRI Segmentation
by: Guo, Bangwei, et al.
Published: (2024)
by: Guo, Bangwei, et al.
Published: (2024)
Similar Items
-
Comparing Zealous and Restrained AI Recommendations in a Real-World Human-AI Collaboration Task
by: Xu, Chengyuan, et al.
Published: (2024) -
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
by: Duan, Lin, et al.
Published: (2025) -
Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AI
by: Xu, Chengyuan, et al.
Published: (2024) -
ScreenAgent: A Vision Language Model-driven Computer Control Agent
by: Niu, Runliang, et al.
Published: (2024) -
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
by: Ouyang, Mingyu, et al.
Published: (2026)