GUI Agents with Reinforcement Learning: Toward Digital Inhabitants
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Junan, Liu, Jian, Lai, Jingxiang, Hu, Jiarui, Sheng, Yiwei, Chen, Shuang, Li, Jian, Du, Dazhao, Guo, Song |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
von: Du, Dazhao, et al.
Veröffentlicht: (2026)
von: Du, Dazhao, et al.
Veröffentlicht: (2026)
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
von: Du, Dazhao, et al.
Veröffentlicht: (2026)
von: Du, Dazhao, et al.
Veröffentlicht: (2026)
ShowUI-Aloha: Human-Taught GUI Agent
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
3D Generation for Embodied AI and Robotic Simulation: A Survey
von: Ye, Tianwei, et al.
Veröffentlicht: (2026)
von: Ye, Tianwei, et al.
Veröffentlicht: (2026)
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
von: Gu, Bohai, et al.
Veröffentlicht: (2026)
von: Gu, Bohai, et al.
Veröffentlicht: (2026)
UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents
von: Xiao, Han, et al.
Veröffentlicht: (2026)
von: Xiao, Han, et al.
Veröffentlicht: (2026)
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
von: Gu, Bohai, et al.
Veröffentlicht: (2026)
von: Gu, Bohai, et al.
Veröffentlicht: (2026)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
von: Xie, Yiping, et al.
Veröffentlicht: (2026)
Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation
von: Gao, Xinzge, et al.
Veröffentlicht: (2025)
von: Gao, Xinzge, et al.
Veröffentlicht: (2025)
Progressive Pretext Task Learning for Human Trajectory Prediction
von: Lin, Xiaotong, et al.
Veröffentlicht: (2024)
von: Lin, Xiaotong, et al.
Veröffentlicht: (2024)
LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
von: Tang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Tang, Jiaqi, et al.
Veröffentlicht: (2025)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
von: Li, Junlong, et al.
Veröffentlicht: (2026)
von: Li, Junlong, et al.
Veröffentlicht: (2026)
Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation
von: Zhou, Zhen, et al.
Veröffentlicht: (2026)
von: Zhou, Zhen, et al.
Veröffentlicht: (2026)
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
Divergence is Uncertainty: A Closed-Form Posterior Covariance for Flow Matching
von: Xing, Jiarui, et al.
Veröffentlicht: (2026)
von: Xing, Jiarui, et al.
Veröffentlicht: (2026)
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
von: Lian, Shuquan, et al.
Veröffentlicht: (2025)
von: Lian, Shuquan, et al.
Veröffentlicht: (2025)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis
von: Wang, Xuan, et al.
Veröffentlicht: (2026)
von: Wang, Xuan, et al.
Veröffentlicht: (2026)
Continual GUI Agents
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
HiMat: DiT-based Ultra-High Resolution SVBRDF Generation
von: Wang, Zixiong, et al.
Veröffentlicht: (2025)
von: Wang, Zixiong, et al.
Veröffentlicht: (2025)
Context Consistency Learning via Sentence Removal for Semi-Supervised Video Paragraph Grounding
von: Zhong, Yaokun, et al.
Veröffentlicht: (2025)
von: Zhong, Yaokun, et al.
Veröffentlicht: (2025)
Benchmarking and Improving GUI Agents in High-Dynamic Environments
von: Liu, Enqi, et al.
Veröffentlicht: (2026)
von: Liu, Enqi, et al.
Veröffentlicht: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
von: Bu, Tianpeng, et al.
Veröffentlicht: (2026)
von: Bu, Tianpeng, et al.
Veröffentlicht: (2026)
AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
von: Shi, Yibo, et al.
Veröffentlicht: (2026)
von: Shi, Yibo, et al.
Veröffentlicht: (2026)
VkSplat: High-Performance 3DGS Training in Vulkan Compute
von: Chen, Jingxiang, et al.
Veröffentlicht: (2026)
von: Chen, Jingxiang, et al.
Veröffentlicht: (2026)
HiconAgent: History Context-aware Policy Optimization for GUI Agents
von: Zhou, Xurui, et al.
Veröffentlicht: (2025)
von: Zhou, Xurui, et al.
Veröffentlicht: (2025)
Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
von: Jiang, Haichao, et al.
Veröffentlicht: (2026)
von: Jiang, Haichao, et al.
Veröffentlicht: (2026)
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
von: Shao, Rui, et al.
Veröffentlicht: (2026)
von: Shao, Rui, et al.
Veröffentlicht: (2026)
A Physical Model-Guided Framework for Underwater Image Enhancement and Depth Estimation
von: Du, Dazhao, et al.
Veröffentlicht: (2024)
von: Du, Dazhao, et al.
Veröffentlicht: (2024)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
Residual-Conditioned Optimal Transport: Towards Structure-Preserving Unpaired and Paired Image Restoration
von: Tang, Xiaole, et al.
Veröffentlicht: (2024)
von: Tang, Xiaole, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
von: Du, Dazhao, et al.
Veröffentlicht: (2026) -
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025) -
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
von: Du, Dazhao, et al.
Veröffentlicht: (2026) -
ShowUI-Aloha: Human-Taught GUI Agent
von: Zhang, Yichun, et al.
Veröffentlicht: (2026) -
3D Generation for Embodied AI and Robotic Simulation: A Survey
von: Ye, Tianwei, et al.
Veröffentlicht: (2026)