InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Ziyun, Wang, Zezhou, Zhang, Xiaoyi, Guo, Zongyu, Li, Jiahao, Li, Bin, Lu, Yan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
by: Zhang, Xiaoyi, et al.
Published: (2025)
by: Zhang, Xiaoyi, et al.
Published: (2025)
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
by: Wang, Junyang, et al.
Published: (2026)
by: Wang, Junyang, et al.
Published: (2026)
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
by: Bu, Tianpeng, et al.
Published: (2026)
by: Bu, Tianpeng, et al.
Published: (2026)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Generative Latent Video Compression
by: Guo, Zongyu, et al.
Published: (2025)
by: Guo, Zongyu, et al.
Published: (2025)
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
by: Tang, Fei, et al.
Published: (2026)
by: Tang, Fei, et al.
Published: (2026)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
by: Wang, Xuehui, et al.
Published: (2025)
by: Wang, Xuehui, et al.
Published: (2025)
CoD: A Diffusion Foundation Model for Image Compression
by: Jia, Zhaoyang, et al.
Published: (2025)
by: Jia, Zhaoyang, et al.
Published: (2025)
Web World Models
by: Feng, Jichen, et al.
Published: (2025)
by: Feng, Jichen, et al.
Published: (2025)
CoD-Lite: Real-Time Diffusion-Based Generative Image Compression
by: Jia, Zhaoyang, et al.
Published: (2026)
by: Jia, Zhaoyang, et al.
Published: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
by: Bai, Hao, et al.
Published: (2026)
by: Bai, Hao, et al.
Published: (2026)
WebGuard: Building a Generalizable Guardrail for Web Agents
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
Generative Video Compression with One-Dimensional Latent Representation
by: Zheng, Zihan, et al.
Published: (2026)
by: Zheng, Zihan, et al.
Published: (2026)
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
by: Zhang, Bofei, et al.
Published: (2025)
by: Zhang, Bofei, et al.
Published: (2025)
Mem-W: Latent Memory-Native GUI Agents
by: Zhang, Guibin, et al.
Published: (2026)
by: Zhang, Guibin, et al.
Published: (2026)
An Efficient Streaming Video Understanding Framework with Agentic Control
by: Liu, Jinming, et al.
Published: (2026)
by: Liu, Jinming, et al.
Published: (2026)
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
by: Zhou, Yuqi, et al.
Published: (2025)
by: Zhou, Yuqi, et al.
Published: (2025)
WebInject: Prompt Injection Attack to Web Agents
by: Wang, Xilong, et al.
Published: (2025)
by: Wang, Xilong, et al.
Published: (2025)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
by: Jia, Yiming, et al.
Published: (2025)
by: Jia, Yiming, et al.
Published: (2025)
Paper2Web: Let's Make Your Paper Alive!
by: Chen, Yuhang, et al.
Published: (2025)
by: Chen, Yuhang, et al.
Published: (2025)
City-on-Web: Real-time Neural Rendering of Large-scale Scenes on the Web
by: Song, Kaiwen, et al.
Published: (2023)
by: Song, Kaiwen, et al.
Published: (2023)
PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
A Survey on (M)LLM-Based GUI Agents
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
by: Zeng, Ziyun, et al.
Published: (2026)
by: Zeng, Ziyun, et al.
Published: (2026)
Benchmarking and Improving GUI Agents in High-Dynamic Environments
by: Liu, Enqi, et al.
Published: (2026)
by: Liu, Enqi, et al.
Published: (2026)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
by: Qin, Yujia, et al.
Published: (2025)
by: Qin, Yujia, et al.
Published: (2025)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
by: Shao, Rui, et al.
Published: (2026)
by: Shao, Rui, et al.
Published: (2026)
FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
by: Lu, Zimu, et al.
Published: (2026)
by: Lu, Zimu, et al.
Published: (2026)
What If We Recaption Billions of Web Images with LLaMA-3?
by: Li, Xianhang, et al.
Published: (2024)
by: Li, Xianhang, et al.
Published: (2024)
IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web
by: Guo, Hongcheng, et al.
Published: (2024)
by: Guo, Hongcheng, et al.
Published: (2024)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
Similar Items
-
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
by: Liu, Xinyi, et al.
Published: (2025) -
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
by: Zhang, Xiaoyi, et al.
Published: (2025) -
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
by: Wang, Junyang, et al.
Published: (2026) -
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
by: Bu, Tianpeng, et al.
Published: (2026) -
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)