UItron: Foundational GUI Agent with Advanced Perception and Planning
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Zhixiong, Huang, Jing, Zheng, Liming, Han, Wenkang, Zhong, Yufeng, Chen, Lei, Yang, Longrong, Chu, Yingjie, He, Yuzhi, Ma, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ScaleTrack: Scaling and back-tracking Automated GUI Agents
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
by: Han, Wenkang, et al.
Published: (2025)
by: Han, Wenkang, et al.
Published: (2025)
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
by: Yang, Longrong, et al.
Published: (2025)
by: Yang, Longrong, et al.
Published: (2025)
DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
MobileDreamer: Generative Sketch World Model for GUI Agent
by: Cao, Yilin, et al.
Published: (2026)
by: Cao, Yilin, et al.
Published: (2026)
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
by: Zhong, Yufeng, et al.
Published: (2026)
by: Zhong, Yufeng, et al.
Published: (2026)
Agentic Reward Modeling: Verifying GUI Agent via Online Proactive Interaction
by: Cui, Chaoqun, et al.
Published: (2026)
by: Cui, Chaoqun, et al.
Published: (2026)
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
by: Chen, Lei, et al.
Published: (2025)
by: Chen, Lei, et al.
Published: (2025)
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
by: Chen, Lei, et al.
Published: (2025)
by: Chen, Lei, et al.
Published: (2025)
TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
by: Jiang, Deyang, et al.
Published: (2026)
by: Jiang, Deyang, et al.
Published: (2026)
RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
by: Zhao, Xuanle, et al.
Published: (2025)
by: Zhao, Xuanle, et al.
Published: (2025)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
by: Chen, Ruihan, et al.
Published: (2025)
by: Chen, Ruihan, et al.
Published: (2025)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
by: Jing, Hongyi, et al.
Published: (2025)
by: Jing, Hongyi, et al.
Published: (2025)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
by: Ye, Xianhang, et al.
Published: (2025)
by: Ye, Xianhang, et al.
Published: (2025)
GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
GUI Agents with Foundation Models: A Comprehensive Survey
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
SpiritSight Agent: Advanced GUI Agent with One Look
by: Huang, Zhiyuan, et al.
Published: (2025)
by: Huang, Zhiyuan, et al.
Published: (2025)
BEAP-Agent: Backtrackable Execution and Adaptive Planning for GUI Agents
by: Lu, Ziyu, et al.
Published: (2026)
by: Lu, Ziyu, et al.
Published: (2026)
CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning
by: Yao, Zhenquan, et al.
Published: (2026)
by: Yao, Zhenquan, et al.
Published: (2026)
scpFormer: A Foundation Model for Unified Representation and Integration of the Single-Cell Proteomics
by: Zhou, Qifeng, et al.
Published: (2026)
by: Zhou, Qifeng, et al.
Published: (2026)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
by: Tang, Liujian, et al.
Published: (2025)
by: Tang, Liujian, et al.
Published: (2025)
A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning
by: Li, Jiahao, et al.
Published: (2025)
by: Li, Jiahao, et al.
Published: (2025)
Boosting Robotic Manipulation Generalization with Minimal Costly Data
by: Zheng, Liming, et al.
Published: (2025)
by: Zheng, Liming, et al.
Published: (2025)
M$^2$-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining
by: Lv, Rui, et al.
Published: (2026)
by: Lv, Rui, et al.
Published: (2026)
See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch
by: Zhang, Xingyi, et al.
Published: (2026)
by: Zhang, Xingyi, et al.
Published: (2026)
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026)
by: Henry, Felix, et al.
Published: (2026)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
by: Zhang, Miaosen, et al.
Published: (2025)
by: Zhang, Miaosen, et al.
Published: (2025)
Optimized Flow Control based on Automatic Differentiation in Compressible Turbulent Channel Flows
by: Wang, Wenkang, et al.
Published: (2024)
by: Wang, Wenkang, et al.
Published: (2024)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
Adaptive Milestone Reward for GUI Agents
by: Zheng, Congmin, et al.
Published: (2026)
by: Zheng, Congmin, et al.
Published: (2026)
GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
by: Zhao, Kangjia, et al.
Published: (2024)
by: Zhao, Kangjia, et al.
Published: (2024)
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
by: Yan, Feng, et al.
Published: (2024)
by: Yan, Feng, et al.
Published: (2024)
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
by: Lian, Shuquan, et al.
Published: (2025)
by: Lian, Shuquan, et al.
Published: (2025)
API Agents vs. GUI Agents: Divergence and Convergence
by: Zhang, Chaoyun, et al.
Published: (2025)
by: Zhang, Chaoyun, et al.
Published: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
by: Shi, Yucheng, et al.
Published: (2025)
by: Shi, Yucheng, et al.
Published: (2025)
Similar Items
-
ScaleTrack: Scaling and back-tracking Automated GUI Agents
by: Huang, Jing, et al.
Published: (2025) -
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
by: Han, Wenkang, et al.
Published: (2025) -
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
by: Yang, Longrong, et al.
Published: (2025) -
DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
by: Zhong, Yufeng, et al.
Published: (2025) -
MobileDreamer: Generative Sketch World Model for GUI Agent
by: Cao, Yilin, et al.
Published: (2026)