WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Hao, Taymanov, Alexey, Zhang, Tong, Kumar, Aviral, Whitehead, Spencer |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
von: Kar, Oğuzhan Fatih, et al.
Veröffentlicht: (2026)
von: Kar, Oğuzhan Fatih, et al.
Veröffentlicht: (2026)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
von: Yang, Rui, et al.
Veröffentlicht: (2026)
von: Yang, Rui, et al.
Veröffentlicht: (2026)
OceanGym: A Benchmark Environment for Underwater Embodied Agents
von: Xue, Yida, et al.
Veröffentlicht: (2025)
von: Xue, Yida, et al.
Veröffentlicht: (2025)
WALT: Web Agents that Learn Tools
von: Prabhu, Viraj, et al.
Veröffentlicht: (2025)
von: Prabhu, Viraj, et al.
Veröffentlicht: (2025)
WebSerial Vision Training for Microcontrollers: A Browser-Based Companion to On-Device CNN Training
von: Ellis, Jeremy
Veröffentlicht: (2026)
von: Ellis, Jeremy
Veröffentlicht: (2026)
Conditional Diffusion on Web-Scale Image Pairs leads to Diverse Image Variations
von: Kumar, Manoj, et al.
Veröffentlicht: (2024)
von: Kumar, Manoj, et al.
Veröffentlicht: (2024)
LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models
von: Li, Fanfei, et al.
Veröffentlicht: (2025)
von: Li, Fanfei, et al.
Veröffentlicht: (2025)
WebInject: Prompt Injection Attack to Web Agents
von: Wang, Xilong, et al.
Veröffentlicht: (2025)
von: Wang, Xilong, et al.
Veröffentlicht: (2025)
Web-based Melanoma Detection
von: Kim, SangHyuk, et al.
Veröffentlicht: (2024)
von: Kim, SangHyuk, et al.
Veröffentlicht: (2024)
No Training Wheels: Steering Vectors for Bias Correction at Inference Time
von: Gupta, Aviral, et al.
Veröffentlicht: (2025)
von: Gupta, Aviral, et al.
Veröffentlicht: (2025)
The Unified Balance Theory of Second-Moment Exponential Scaling Optimizers in Visual Tasks
von: Zhang, Gongyue, et al.
Veröffentlicht: (2024)
von: Zhang, Gongyue, et al.
Veröffentlicht: (2024)
Walking the Web of Concept-Class Relationships in Incrementally Trained Interpretable Models
von: Agrawal, Susmit, et al.
Veröffentlicht: (2025)
von: Agrawal, Susmit, et al.
Veröffentlicht: (2025)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
von: Zhang, Ziyun, et al.
Veröffentlicht: (2026)
von: Zhang, Ziyun, et al.
Veröffentlicht: (2026)
HistoGym: A Reinforcement Learning Environment for Histopathological Image Analysis
von: Liu, Zhi-Bo, et al.
Veröffentlicht: (2024)
von: Liu, Zhi-Bo, et al.
Veröffentlicht: (2024)
Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions
von: Chen, Hongyi, et al.
Veröffentlicht: (2025)
von: Chen, Hongyi, et al.
Veröffentlicht: (2025)
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
von: Cheang, Chi-Lam, et al.
Veröffentlicht: (2024)
von: Cheang, Chi-Lam, et al.
Veröffentlicht: (2024)
EEO-TFV: Escape-Explore Optimizer for Web-Scale Time-Series Forecasting and Vision Analysis
von: Wang, Hua, et al.
Veröffentlicht: (2026)
von: Wang, Hua, et al.
Veröffentlicht: (2026)
Noise-Tolerant Hybrid Prototypical Learning with Noisy Web Data
von: Liang, Chao, et al.
Veröffentlicht: (2025)
von: Liang, Chao, et al.
Veröffentlicht: (2025)
Structuring a Training Strategy to Robustify Perception Models with Realistic Image Augmentations
von: Hammam, Ahmed, et al.
Veröffentlicht: (2024)
von: Hammam, Ahmed, et al.
Veröffentlicht: (2024)
TimeWarp: Evaluating Web Agents by Revisiting the Past
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2026)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2026)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
von: Gupta, Tanmay, et al.
Veröffentlicht: (2026)
von: Gupta, Tanmay, et al.
Veröffentlicht: (2026)
DRAGON: A Large-Scale Dataset of Realistic Images Generated by Diffusion Models
von: Bertazzini, Giulia, et al.
Veröffentlicht: (2025)
von: Bertazzini, Giulia, et al.
Veröffentlicht: (2025)
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
von: Tang, Jingqun, et al.
Veröffentlicht: (2024)
von: Tang, Jingqun, et al.
Veröffentlicht: (2024)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
von: Wang, Bowen, et al.
Veröffentlicht: (2026)
von: Wang, Bowen, et al.
Veröffentlicht: (2026)
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
von: Chen, William, et al.
Veröffentlicht: (2024)
von: Chen, William, et al.
Veröffentlicht: (2024)
InnoGym: Benchmarking the Innovation Potential of AI Agents
von: Zhang, Jintian, et al.
Veröffentlicht: (2025)
von: Zhang, Jintian, et al.
Veröffentlicht: (2025)
Descripción automática de secciones delgadas de rocas: una aplicación Web
von: Paucar, Stalyn, et al.
Veröffentlicht: (2024)
von: Paucar, Stalyn, et al.
Veröffentlicht: (2024)
Growing Visual Generative Capacity for Pre-Trained MLLMs
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
von: He, Haoran, et al.
Veröffentlicht: (2024)
von: He, Haoran, et al.
Veröffentlicht: (2024)
Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos
von: Shi, Junyao, et al.
Veröffentlicht: (2025)
von: Shi, Junyao, et al.
Veröffentlicht: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
von: Mercea, Otniel-Bogdan, et al.
Veröffentlicht: (2024)
Visual Test-time Scaling for GUI Agent Grounding
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
Hierarchical Invariance for Robust and Interpretable Vision Tasks at Larger Scales
von: Qi, Shuren, et al.
Veröffentlicht: (2024)
von: Qi, Shuren, et al.
Veröffentlicht: (2024)
Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream
von: Gokce, Abdulkadir, et al.
Veröffentlicht: (2024)
von: Gokce, Abdulkadir, et al.
Veröffentlicht: (2024)
Generating a Biometrically Unique and Realistic Iris Database
von: Zhang, Jingxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jingxuan, et al.
Veröffentlicht: (2025)
WebCryptoAgent: Agentic Crypto Trading with Web Informatics
von: Kurban, Ali, et al.
Veröffentlicht: (2026)
von: Kurban, Ali, et al.
Veröffentlicht: (2026)
Tensor-Train Point Cloud Compression and Efficient Approximate Nearest-Neighbor Search
von: Novikov, Georgii, et al.
Veröffentlicht: (2024)
von: Novikov, Georgii, et al.
Veröffentlicht: (2024)
Vision Learners Meet Web Image-Text Pairs
von: Zhao, Bingchen, et al.
Veröffentlicht: (2023)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024) -
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
von: Kar, Oğuzhan Fatih, et al.
Veröffentlicht: (2026) -
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
von: Yang, Rui, et al.
Veröffentlicht: (2026) -
OceanGym: A Benchmark Environment for Underwater Embodied Agents
von: Xue, Yida, et al.
Veröffentlicht: (2025) -
WALT: Web Agents that Learn Tools
von: Prabhu, Viraj, et al.
Veröffentlicht: (2025)