VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Yunpeng, Bian, Yiheng, Tang, Yongtao, Ma, Guiyu, Cai, Zhongmin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information Seeking
by: Bian, Yiheng, et al.
Published: (2025)
by: Bian, Yiheng, et al.
Published: (2025)
Predicting User Behavior in Smart Spaces with LLM-Enhanced Logs and Personalized Prompts
by: Song, Yunpeng, et al.
Published: (2024)
by: Song, Yunpeng, et al.
Published: (2024)
GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
by: Vu, Minh Duc, et al.
Published: (2024)
by: Vu, Minh Duc, et al.
Published: (2024)
Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
by: Huang, Tian, et al.
Published: (2024)
by: Huang, Tian, et al.
Published: (2024)
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
by: Zhang, Li, et al.
Published: (2024)
by: Zhang, Li, et al.
Published: (2024)
Context-Aware Workflow Decomposition for Automated Mobile UI Annotation Using Multimodal Large Language Models
by: Parvez, Athar, et al.
Published: (2026)
by: Parvez, Athar, et al.
Published: (2026)
MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding
by: Parvez, Athar, et al.
Published: (2026)
by: Parvez, Athar, et al.
Published: (2026)
MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
by: Kong, Yi, et al.
Published: (2025)
by: Kong, Yi, et al.
Published: (2025)
DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
by: Xu, Yuan, et al.
Published: (2025)
by: Xu, Yuan, et al.
Published: (2025)
CrowdGenUI: Aligning LLM-Based UI Generation with Crowdsourced User Preferences
by: Liu, Yimeng, et al.
Published: (2024)
by: Liu, Yimeng, et al.
Published: (2024)
In-Situ Mode: Generative AI-Driven Characters Transforming Art Engagement Through Anthropomorphic Narratives
by: Li, Yongming, et al.
Published: (2024)
by: Li, Yongming, et al.
Published: (2024)
CAAP: Context-Aware Action Planning Prompting to Solve Computer Tasks with Front-End UI Only
by: Cho, Junhee, et al.
Published: (2024)
by: Cho, Junhee, et al.
Published: (2024)
GhostUI: Unveiling Hidden Interactions in Mobile UI
by: Kweon, Minkyu, et al.
Published: (2026)
by: Kweon, Minkyu, et al.
Published: (2026)
LightVA: Lightweight Visual Analytics with LLM Agent-Based Task Planning and Execution
by: Zhao, Yuheng, et al.
Published: (2024)
by: Zhao, Yuheng, et al.
Published: (2024)
MLLM-Based UI2Code Automation Guided by UI Layout Information
by: Wu, Fan, et al.
Published: (2025)
by: Wu, Fan, et al.
Published: (2025)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024)
by: You, Keen, et al.
Published: (2024)
Agent-Initiated Interaction in Phone UI Automation
by: Kahlon, Noam, et al.
Published: (2025)
by: Kahlon, Noam, et al.
Published: (2025)
User-Centric Design of UI for Mobile Banking Apps: Improving UI and Features for Better Customer Experience
by: Chitrakar, Luniva, et al.
Published: (2026)
by: Chitrakar, Luniva, et al.
Published: (2026)
Simulating Vision Impairment in Virtual Reality -- A Comparison of Visual Task Performance with Real and Simulated Tunnel Vision
by: Neugebauer, Alexander, et al.
Published: (2023)
by: Neugebauer, Alexander, et al.
Published: (2023)
Automated UI Interface Generation via Diffusion Models: Enhancing Personalization and Efficiency
by: Duan, Yifei, et al.
Published: (2025)
by: Duan, Yifei, et al.
Published: (2025)
From Interaction to Impact: Towards Safer AI Agents Through Understanding and Evaluating Mobile UI Operation Impacts
by: Zhang, Zhuohao Jerry, et al.
Published: (2024)
by: Zhang, Zhuohao Jerry, et al.
Published: (2024)
Macaron-A2UI: A Model for Generative UI in Personal Agents
by: Kong, Fancy, et al.
Published: (2026)
by: Kong, Fancy, et al.
Published: (2026)
ProactiveVA: Proactive Visual Analytics with LLM-Based UI Agent
by: Zhao, Yuheng, et al.
Published: (2025)
by: Zhao, Yuheng, et al.
Published: (2025)
Automating UI Optimization through Multi-Agentic Reasoning
by: Li, Zhipeng, et al.
Published: (2026)
by: Li, Zhipeng, et al.
Published: (2026)
MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task Automation
by: Zhu, Zichen, et al.
Published: (2024)
by: Zhu, Zichen, et al.
Published: (2024)
MarkupLens: Balancing Computer Vision Assistance and Control in Professional Video Annotation for Video-Based Design Tasks
by: He, Tianhao, et al.
Published: (2024)
by: He, Tianhao, et al.
Published: (2024)
Chartist: Task-driven Eye Movement Control for Chart Reading
by: Shi, Danqing, et al.
Published: (2025)
by: Shi, Danqing, et al.
Published: (2025)
From Following to Understanding: Investigating the Role of Reflective Prompts in AR-Guided Tasks to Promote Task Understanding
by: Zhang, Nandi, et al.
Published: (2025)
by: Zhang, Nandi, et al.
Published: (2025)
Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision Models
by: Zhao, Ada Yi, et al.
Published: (2025)
by: Zhao, Ada Yi, et al.
Published: (2025)
CogInstrument: Modeling Cognitive Processes for Bidirectional Human-LLM Alignment in Planning Tasks
by: Wang, Anqi, et al.
Published: (2026)
by: Wang, Anqi, et al.
Published: (2026)
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
by: Schoop, Eldon, et al.
Published: (2022)
by: Schoop, Eldon, et al.
Published: (2022)
Explore, Select, Derive, and Recall: Augmenting LLM with Human-like Memory for Mobile Task Automation
by: Lee, Sunjae, et al.
Published: (2023)
by: Lee, Sunjae, et al.
Published: (2023)
GraphPilot: GUI Task Automation with One-Step LLM Reasoning Powered by Knowledge Graph
by: Yu, Mingxian, et al.
Published: (2026)
by: Yu, Mingxian, et al.
Published: (2026)
JumpStarter: Human-AI Planning with Task-Structured Context Curation
by: Zhang, Xuanming, et al.
Published: (2024)
by: Zhang, Xuanming, et al.
Published: (2024)
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
by: Feng, Sidong, et al.
Published: (2024)
by: Feng, Sidong, et al.
Published: (2024)
RWKV-UI: UI Understanding with Enhanced Perception and Reasoning
by: Yang, Jiaxi, et al.
Published: (2025)
by: Yang, Jiaxi, et al.
Published: (2025)
Understanding the Human-LLM Dynamic: A Literature Survey of LLM Use in Programming Tasks
by: Etsenake, Deborah, et al.
Published: (2024)
by: Etsenake, Deborah, et al.
Published: (2024)
Preference-Guided Multi-Objective UI Adaptation
by: Song, Yao, et al.
Published: (2025)
by: Song, Yao, et al.
Published: (2025)
Game Master LLM: Task-Based Role-Playing for Natural Slang Learning
by: Tahmasbi, Amir, et al.
Published: (2025)
by: Tahmasbi, Amir, et al.
Published: (2025)
Task-Aware Delegation Cues for LLM Agents
by: Gu, Xingrui
Published: (2026)
by: Gu, Xingrui
Published: (2026)
Similar Items
-
DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information Seeking
by: Bian, Yiheng, et al.
Published: (2025) -
Predicting User Behavior in Smart Spaces with LLM-Enhanced Logs and Personalized Prompts
by: Song, Yunpeng, et al.
Published: (2024) -
GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
by: Vu, Minh Duc, et al.
Published: (2024) -
Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
by: Huang, Tian, et al.
Published: (2024) -
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
by: Zhang, Li, et al.
Published: (2024)