AnyUser: Translating Sketched User Intent into Domestic Robots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Songyuan, Tan, Huibin, Yang, Kailun, Yang, Wenjing, Yang, Shaowu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915917925974016
author Yang, Songyuan
Tan, Huibin
Yang, Kailun
Yang, Wenjing
Yang, Shaowu
author_facet Yang, Songyuan
Tan, Huibin
Yang, Kailun
Yang, Wenjing
Yang, Shaowu
contents We introduce AnyUser, a unified robotic instruction system for intuitive domestic task instruction via free-form sketches on camera images, optionally with language. AnyUser interprets multimodal inputs (sketch, vision, language) as spatial-semantic primitives to generate executable robot actions requiring no prior maps or models. Novel components include multimodal fusion for understanding and a hierarchical policy for robust action generation. Efficacy is shown via extensive evaluations: (1) Quantitative benchmarks on the large-scale dataset showing high accuracy in interpreting diverse sketch-based commands across various simulated domestic scenes. (2) Real-world validation on two distinct robotic platforms, a statically mounted 7-DoF assistive arm (KUKA LBR iiwa) and a dual-arm mobile manipulator (Realman RMC-AIDAL), performing representative tasks like targeted wiping and area cleaning, confirming the system's ability to ground instructions and execute them reliably in physical environments. (3) A comprehensive user study involving diverse demographics (elderly, simulated non-verbal, low technical literacy) demonstrating significant improvements in usability and task specification efficiency, achieving high task completion rates (85.7%-96.4%) and user satisfaction. AnyUser bridges the gap between advanced robotic capabilities and the need for accessible non-expert interaction, laying the foundation for practical assistive robots adaptable to real-world human environments.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04811
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AnyUser: Translating Sketched User Intent into Domestic Robots
Yang, Songyuan
Tan, Huibin
Yang, Kailun
Yang, Wenjing
Yang, Shaowu
Robotics
Computer Vision and Pattern Recognition
Human-Computer Interaction
We introduce AnyUser, a unified robotic instruction system for intuitive domestic task instruction via free-form sketches on camera images, optionally with language. AnyUser interprets multimodal inputs (sketch, vision, language) as spatial-semantic primitives to generate executable robot actions requiring no prior maps or models. Novel components include multimodal fusion for understanding and a hierarchical policy for robust action generation. Efficacy is shown via extensive evaluations: (1) Quantitative benchmarks on the large-scale dataset showing high accuracy in interpreting diverse sketch-based commands across various simulated domestic scenes. (2) Real-world validation on two distinct robotic platforms, a statically mounted 7-DoF assistive arm (KUKA LBR iiwa) and a dual-arm mobile manipulator (Realman RMC-AIDAL), performing representative tasks like targeted wiping and area cleaning, confirming the system's ability to ground instructions and execute them reliably in physical environments. (3) A comprehensive user study involving diverse demographics (elderly, simulated non-verbal, low technical literacy) demonstrating significant improvements in usability and task specification efficiency, achieving high task completion rates (85.7%-96.4%) and user satisfaction. AnyUser bridges the gap between advanced robotic capabilities and the need for accessible non-expert interaction, laying the foundation for practical assistive robots adaptable to real-world human environments.
title AnyUser: Translating Sketched User Intent into Domestic Robots
topic Robotics
Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2604.04811