LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Guangyi, Zhao, Pengxiang, Liu, Liang, Chen, Zhiming, Chai, Yuxiang, Ren, Shuai, Wang, Hao, He, Shibo, Meng, Wenchao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
by: Liu, Guangyi, et al.
Published: (2025)
by: Liu, Guangyi, et al.
Published: (2025)
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
by: Chai, Yuxiang, et al.
Published: (2024)
by: Chai, Yuxiang, et al.
Published: (2024)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
by: Sun, Jiazheng, et al.
Published: (2026)
by: Sun, Jiazheng, et al.
Published: (2026)
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
by: Gong, Yichen, et al.
Published: (2026)
by: Gong, Yichen, et al.
Published: (2026)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
by: Tang, Liujian, et al.
Published: (2025)
by: Tang, Liujian, et al.
Published: (2025)
Design and Evaluation of an AI-DrivenPersonalized Mobile App to Provide MultifacetedHealth Support for Type 2 Diabetes Patients inChina
by: Meng, Yibo, et al.
Published: (2025)
by: Meng, Yibo, et al.
Published: (2025)
GUI Agents with Foundation Models: A Comprehensive Survey
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
MobileViews: A Million-scale and Diverse Mobile GUI Dataset
by: Gao, Longxi, et al.
Published: (2024)
by: Gao, Longxi, et al.
Published: (2024)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026)
by: Henry, Felix, et al.
Published: (2026)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
by: Wu, Zongru, et al.
Published: (2025)
by: Wu, Zongru, et al.
Published: (2025)
MERBench: A Unified Evaluation Benchmark for Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2024)
by: Lian, Zheng, et al.
Published: (2024)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
SummAct: Uncovering User Intentions Through Interactive Behaviour Summarisation
by: Zhang, Guanhua, et al.
Published: (2024)
by: Zhang, Guanhua, et al.
Published: (2024)
Beyond Chat and Clicks: GUI Agents for In-Situ Assistance via Live Interface Transformation
by: Hao, Pan, et al.
Published: (2026)
by: Hao, Pan, et al.
Published: (2026)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025)
by: Nong, Songqin, et al.
Published: (2025)
AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
by: Kim, Jeonghyeon, et al.
Published: (2026)
by: Kim, Jeonghyeon, et al.
Published: (2026)
GUI Agents: A Survey
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
by: Wang, Luyuan, et al.
Published: (2024)
by: Wang, Luyuan, et al.
Published: (2024)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
by: Lee, Jungjae, et al.
Published: (2025)
by: Lee, Jungjae, et al.
Published: (2025)
Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing
by: Zhang, Shuning, et al.
Published: (2025)
by: Zhang, Shuning, et al.
Published: (2025)
EZBlender: Efficient 3D Editing with Plan-and-ReAct Agent
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
API Agents vs. GUI Agents: Divergence and Convergence
by: Zhang, Chaoyun, et al.
Published: (2025)
by: Zhang, Chaoyun, et al.
Published: (2025)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
by: Gebreegziabher, Simret Araya, et al.
Published: (2026)
by: Gebreegziabher, Simret Araya, et al.
Published: (2026)
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
by: Zhao, Pengxiang, et al.
Published: (2025)
by: Zhao, Pengxiang, et al.
Published: (2025)
Demonstrating HumanTHOR: A Simulation Platform and Benchmark for Human-Robot Collaboration in a Shared Workspace
by: Wang, Chenxu, et al.
Published: (2024)
by: Wang, Chenxu, et al.
Published: (2024)
TinyClick: Single-Turn Agent for Empowering GUI Automation
by: Pawlowski, Pawel, et al.
Published: (2024)
by: Pawlowski, Pawel, et al.
Published: (2024)
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
by: Hu, Haobo, et al.
Published: (2026)
by: Hu, Haobo, et al.
Published: (2026)
Final Happiness: What Intelligent User Interfaces Can Do for The Lonely Dying
by: Meng, Yibo, et al.
Published: (2025)
by: Meng, Yibo, et al.
Published: (2025)
Large Language Model-Brained GUI Agents: A Survey
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
AI Chatbots as Professional Service Agents: Developing a Professional Identity
by: Li, Wenwen, et al.
Published: (2025)
by: Li, Wenwen, et al.
Published: (2025)
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
by: Liu, Guangyi, et al.
Published: (2026)
by: Liu, Guangyi, et al.
Published: (2026)
A Unified, Cross-Platform Framework for Automatic GUI and Plugin Generation in Structural Bioinformatics and Beyond
by: Guo, Sikao, et al.
Published: (2026)
by: Guo, Sikao, et al.
Published: (2026)
Legacy Learning Strategy Based on Few-Shot Font Generation Models for Automatic Text Design in Metaverse Content
by: Kim, Younghwi, et al.
Published: (2024)
by: Kim, Younghwi, et al.
Published: (2024)
Neighbor-Environment Observer: An Intelligent Agent for Immersive Working Companionship
by: Sun, Zhe, et al.
Published: (2024)
by: Sun, Zhe, et al.
Published: (2024)
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
by: Tang, Jingyu, et al.
Published: (2025)
by: Tang, Jingyu, et al.
Published: (2025)
From Human Negotiation to Agent Negotiation: Personal Mobility Agents in Automated Traffic
by: Jansen, Pascal
Published: (2026)
by: Jansen, Pascal
Published: (2026)
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
by: Sun, Qiushi, et al.
Published: (2025)
by: Sun, Qiushi, et al.
Published: (2025)
Think, Act, and Ask: Open-World Interactive Personalized Robot Navigation
by: Dai, Yinpei, et al.
Published: (2023)
by: Dai, Yinpei, et al.
Published: (2023)
Similar Items
-
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
by: Liu, Guangyi, et al.
Published: (2025) -
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
by: Chai, Yuxiang, et al.
Published: (2024) -
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
by: Sun, Jiazheng, et al.
Published: (2026) -
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
by: Gong, Yichen, et al.
Published: (2026) -
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
by: Tang, Liujian, et al.
Published: (2025)