TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Bofei, Shang, Zirui, Gao, Zhi, Zhang, Wang, Xie, Rui, Ma, Xiaojian, Yuan, Tao, Wu, Xinxiao, Zhu, Song-Chun, Li, Qing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
di: Xie, Rui, et al.
Pubblicazione: (2026)
di: Xie, Rui, et al.
Pubblicazione: (2026)
UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics
di: Wu, Mengzhou, et al.
Pubblicazione: (2026)
di: Wu, Mengzhou, et al.
Pubblicazione: (2026)
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
di: Gao, Zhi, et al.
Pubblicazione: (2024)
di: Gao, Zhi, et al.
Pubblicazione: (2024)
Enhancing Web Agents with a Hierarchical Memory Tree
di: Tan, Yunteng, et al.
Pubblicazione: (2026)
di: Tan, Yunteng, et al.
Pubblicazione: (2026)
LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization
di: Shang, Zirui, et al.
Pubblicazione: (2025)
di: Shang, Zirui, et al.
Pubblicazione: (2025)
ShowUI-Aloha: Human-Taught GUI Agent
di: Zhang, Yichun, et al.
Pubblicazione: (2026)
di: Zhang, Yichun, et al.
Pubblicazione: (2026)
FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models
di: Li, Pengxiang, et al.
Pubblicazione: (2024)
di: Li, Pengxiang, et al.
Pubblicazione: (2024)
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
di: Xu, Yiheng, et al.
Pubblicazione: (2024)
di: Xu, Yiheng, et al.
Pubblicazione: (2024)
Temac: Multi-Agent Collaboration for Automated Web GUI Testing
di: Liu, Chenxu, et al.
Pubblicazione: (2025)
di: Liu, Chenxu, et al.
Pubblicazione: (2025)
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
di: Zhang, Shaojie, et al.
Pubblicazione: (2025)
di: Zhang, Shaojie, et al.
Pubblicazione: (2025)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
di: Qin, Yujia, et al.
Pubblicazione: (2025)
di: Qin, Yujia, et al.
Pubblicazione: (2025)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
di: Xiong, Weimin, et al.
Pubblicazione: (2026)
di: Xiong, Weimin, et al.
Pubblicazione: (2026)
MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
di: Zhou, Hanzhang, et al.
Pubblicazione: (2025)
di: Zhou, Hanzhang, et al.
Pubblicazione: (2025)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
di: Shao, Rui, et al.
Pubblicazione: (2026)
di: Shao, Rui, et al.
Pubblicazione: (2026)
UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents
di: Chai, Yuxiang, et al.
Pubblicazione: (2026)
di: Chai, Yuxiang, et al.
Pubblicazione: (2026)
From Task to Tutorial: An Automated GUI Framework for Excel Tutorial Document and Video Creation
di: Xie, Yuhang, et al.
Pubblicazione: (2025)
di: Xie, Yuhang, et al.
Pubblicazione: (2025)
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents
di: Pahuja, Vardaan, et al.
Pubblicazione: (2025)
di: Pahuja, Vardaan, et al.
Pubblicazione: (2025)
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
di: Yang, Zhen, et al.
Pubblicazione: (2025)
di: Yang, Zhen, et al.
Pubblicazione: (2025)
AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
di: Li, Yuankai, et al.
Pubblicazione: (2026)
di: Li, Yuankai, et al.
Pubblicazione: (2026)
Video Summarization using Denoising Diffusion Probabilistic Model
di: Shang, Zirui, et al.
Pubblicazione: (2024)
di: Shang, Zirui, et al.
Pubblicazione: (2024)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
di: Zhang, Ziyun, et al.
Pubblicazione: (2026)
di: Zhang, Ziyun, et al.
Pubblicazione: (2026)
EmbeWebAgent: Embedding Web Agents into Any Customized UI
di: Ma, Chenyang, et al.
Pubblicazione: (2026)
di: Ma, Chenyang, et al.
Pubblicazione: (2026)
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
di: Lian, Shuquan, et al.
Pubblicazione: (2025)
di: Lian, Shuquan, et al.
Pubblicazione: (2025)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
di: Wang, Haoming, et al.
Pubblicazione: (2025)
di: Wang, Haoming, et al.
Pubblicazione: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2024)
PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
di: Liu, Zikang, et al.
Pubblicazione: (2025)
di: Liu, Zikang, et al.
Pubblicazione: (2025)
Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
di: Lu, Junyu, et al.
Pubblicazione: (2025)
di: Lu, Junyu, et al.
Pubblicazione: (2025)
Aria-UI: Visual Grounding for GUI Instructions
di: Yang, Yuhao, et al.
Pubblicazione: (2024)
di: Yang, Yuhao, et al.
Pubblicazione: (2024)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
di: Tang, Fei, et al.
Pubblicazione: (2026)
di: Tang, Fei, et al.
Pubblicazione: (2026)
From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
di: Ran, Dezhi, et al.
Pubblicazione: (2025)
di: Ran, Dezhi, et al.
Pubblicazione: (2025)
WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis
di: Gao, Yifei, et al.
Pubblicazione: (2025)
di: Gao, Yifei, et al.
Pubblicazione: (2025)
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
di: Liu, Yuhang, et al.
Pubblicazione: (2025)
di: Liu, Yuhang, et al.
Pubblicazione: (2025)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
di: Xiao, Han, et al.
Pubblicazione: (2025)
di: Xiao, Han, et al.
Pubblicazione: (2025)
UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
di: Lu, Zhengxi, et al.
Pubblicazione: (2025)
di: Lu, Zhengxi, et al.
Pubblicazione: (2025)
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
di: Lin, Zichuan, et al.
Pubblicazione: (2026)
di: Lin, Zichuan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
di: Xie, Rui, et al.
Pubblicazione: (2026) -
UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics
di: Wu, Mengzhou, et al.
Pubblicazione: (2026) -
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
di: Li, Pengxiang, et al.
Pubblicazione: (2025) -
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
di: Li, Pengxiang, et al.
Pubblicazione: (2025) -
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
di: Gao, Zhi, et al.
Pubblicazione: (2024)