Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kar, Oğuzhan Fatih, Bachmann, Roman, Gong, Yuanzheng, Larsen, Anders Boesen Lindbo, Dehghan, Afshin
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913099406114816
author Kar, Oğuzhan Fatih
Bachmann, Roman
Gong, Yuanzheng
Larsen, Anders Boesen Lindbo
Dehghan, Afshin
author_facet Kar, Oğuzhan Fatih
Bachmann, Roman
Gong, Yuanzheng
Larsen, Anders Boesen Lindbo
Dehghan, Afshin
contents The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories for supervised fine-tuning or a handful of simulated environments for RL training, thus failing to capture web diversity. We propose Weblica (Web Replica), a framework for constructing reproducible and scalable web environments. Our framework leverages 1) HTTP-level caching to capture and replay stable visual states while preserving interactive behavior and 2) LLM-based environment synthesis grounded in real-world websites and core web navigation skills. Using this framework, we scale RL training to thousands of diverse environments and tasks. Our best model, Weblica-8B, outperforms open-weight baselines of similar size across multiple web navigation benchmarks while using fewer inference steps, scales favorably with additional test-time compute, and is competitive with API models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06761
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
Kar, Oğuzhan Fatih
Bachmann, Roman
Gong, Yuanzheng
Larsen, Anders Boesen Lindbo
Dehghan, Afshin
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories for supervised fine-tuning or a handful of simulated environments for RL training, thus failing to capture web diversity. We propose Weblica (Web Replica), a framework for constructing reproducible and scalable web environments. Our framework leverages 1) HTTP-level caching to capture and replay stable visual states while preserving interactive behavior and 2) LLM-based environment synthesis grounded in real-world websites and core web navigation skills. Using this framework, we scale RL training to thousands of diverse environments and tasks. Our best model, Weblica-8B, outperforms open-weight baselines of similar size across multiple web navigation benchmarks while using fewer inference steps, scales favorably with additional test-time compute, and is competitive with API models.
title Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.06761