META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Liangtai, Chen, Xingyu, Chen, Lu, Dai, Tianle, Zhu, Zichen, Yu, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ProgRM: Build Better GUI Agents with Progress Rewards
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
von: Xu, Haiyang, et al.
Veröffentlicht: (2026)
von: Xu, Haiyang, et al.
Veröffentlicht: (2026)
Adaptive Milestone Reward for GUI Agents
von: Zheng, Congmin, et al.
Veröffentlicht: (2026)
von: Zheng, Congmin, et al.
Veröffentlicht: (2026)
Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
von: Zhang, Junlei, et al.
Veröffentlicht: (2025)
von: Zhang, Junlei, et al.
Veröffentlicht: (2025)
RISK: A Framework for GUI Agents in E-commerce Risk Management
von: Chen, Renqi, et al.
Veröffentlicht: (2025)
von: Chen, Renqi, et al.
Veröffentlicht: (2025)
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
von: Han, Wenkang, et al.
Veröffentlicht: (2025)
von: Han, Wenkang, et al.
Veröffentlicht: (2025)
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
von: Tang, Fei, et al.
Veröffentlicht: (2026)
von: Tang, Fei, et al.
Veröffentlicht: (2026)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
Retrieval-augmented GUI Agents with Generative Guidelines
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
A Survey on (M)LLM-Based GUI Agents
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
von: Yang, Rui, et al.
Veröffentlicht: (2026)
von: Yang, Rui, et al.
Veröffentlicht: (2026)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
von: Guo, Ruoqi, et al.
Veröffentlicht: (2026)
von: Guo, Ruoqi, et al.
Veröffentlicht: (2026)
History-Aware Reasoning for GUI Agents
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
von: Lee, Jungjae, et al.
Veröffentlicht: (2025)
von: Lee, Jungjae, et al.
Veröffentlicht: (2025)
BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2025)
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2025)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents
von: Zhao, Yuan, et al.
Veröffentlicht: (2025)
von: Zhao, Yuan, et al.
Veröffentlicht: (2025)
VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation
von: Han, Qijun, et al.
Veröffentlicht: (2026)
von: Han, Qijun, et al.
Veröffentlicht: (2026)
D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
von: Chen, Sen, et al.
Veröffentlicht: (2025)
von: Chen, Sen, et al.
Veröffentlicht: (2025)
Large Language Model-Brained GUI Agents: A Survey
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
OmniParser for Pure Vision Based GUI Agent
von: Lu, Yadong, et al.
Veröffentlicht: (2024)
von: Lu, Yadong, et al.
Veröffentlicht: (2024)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
von: Lu, Quanfeng, et al.
Veröffentlicht: (2025)
von: Lu, Quanfeng, et al.
Veröffentlicht: (2025)
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2026)
von: Tang, Fei, et al.
Veröffentlicht: (2026)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
von: Tang, Liujian, et al.
Veröffentlicht: (2025)
von: Tang, Liujian, et al.
Veröffentlicht: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
von: Henry, Felix, et al.
Veröffentlicht: (2026)
von: Henry, Felix, et al.
Veröffentlicht: (2026)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
von: Zhang, Ziyun, et al.
Veröffentlicht: (2026)
von: Zhang, Ziyun, et al.
Veröffentlicht: (2026)
Auto-scaling Continuous Memory for GUI Agent
von: Wu, Wenyi, et al.
Veröffentlicht: (2025)
von: Wu, Wenyi, et al.
Veröffentlicht: (2025)
GUICourse: From General Vision Language Models to Versatile GUI Agents
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ProgRM: Build Better GUI Agents with Progress Rewards
von: Zhang, Danyang, et al.
Veröffentlicht: (2025) -
Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
von: Xu, Haiyang, et al.
Veröffentlicht: (2026) -
Adaptive Milestone Reward for GUI Agents
von: Zheng, Congmin, et al.
Veröffentlicht: (2026) -
Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
von: Zhang, Junlei, et al.
Veröffentlicht: (2025) -
RISK: A Framework for GUI Agents in E-commerce Risk Management
von: Chen, Renqi, et al.
Veröffentlicht: (2025)