Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Kar, Oğuzhan Fatih, Bachmann, Roman, Gong, Yuanzheng, Larsen, Anders Boesen Lindbo, Dehghan, Afshin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
by: Bachmann, Roman, et al.
Published: (2024)
by: Bachmann, Roman, et al.
Published: (2024)
(1D) Ordered Tokens Enable Efficient Test-Time Search
by: Gao, Zhitong, et al.
Published: (2026)
by: Gao, Zhitong, et al.
Published: (2026)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025)
by: Ramachandran, Rahul, et al.
Published: (2025)
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
by: Atanov, Andrei, et al.
Published: (2026)
by: Atanov, Andrei, et al.
Published: (2026)
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
by: Bachmann, Roman, et al.
Published: (2025)
by: Bachmann, Roman, et al.
Published: (2025)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
by: Zhang, Ziyun, et al.
Published: (2026)
by: Zhang, Ziyun, et al.
Published: (2026)
WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark
by: Yuan, Peng, et al.
Published: (2026)
by: Yuan, Peng, et al.
Published: (2026)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping
by: Lazarow, Justin, et al.
Published: (2025)
by: Lazarow, Justin, et al.
Published: (2025)
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
by: Bai, Hao, et al.
Published: (2026)
by: Bai, Hao, et al.
Published: (2026)
Learning Generative Interactive Environments By Trained Agent Exploration
by: Kazemi, Naser, et al.
Published: (2024)
by: Kazemi, Naser, et al.
Published: (2024)
AToken: A Unified Tokenizer for Vision
by: Lu, Jiasen, et al.
Published: (2025)
by: Lu, Jiasen, et al.
Published: (2025)
Pic2Diagnosis: A Method for Diagnosis of Cardiovascular Diseases from the Printed ECG Pictures
by: Büyüksolak, Oğuzhan, et al.
Published: (2025)
by: Büyüksolak, Oğuzhan, et al.
Published: (2025)
MEDVISTAGYM: A Scalable Training Environment for Thinking with Medical Images via Tool-Integrated Reinforcement Learning
by: Lu, Meng, et al.
Published: (2026)
by: Lu, Meng, et al.
Published: (2026)
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
by: Wu, Changti, et al.
Published: (2026)
by: Wu, Changti, et al.
Published: (2026)
Cubify Anything: Scaling Indoor 3D Object Detection
by: Lazarow, Justin, et al.
Published: (2024)
by: Lazarow, Justin, et al.
Published: (2024)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
by: Tian, Rui, et al.
Published: (2025)
by: Tian, Rui, et al.
Published: (2025)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
WebSight: A Vision-First Architecture for Robust Web Agents
by: Bhathal, Tanvir, et al.
Published: (2025)
by: Bhathal, Tanvir, et al.
Published: (2025)
WebInject: Prompt Injection Attack to Web Agents
by: Wang, Xilong, et al.
Published: (2025)
by: Wang, Xilong, et al.
Published: (2025)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024)
by: Xu, Mingze, et al.
Published: (2024)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026)
by: Guo, JunJia, et al.
Published: (2026)
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
by: Wang, Haibo, et al.
Published: (2025)
by: Wang, Haibo, et al.
Published: (2025)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
by: Jang, Lawrence, et al.
Published: (2024)
by: Jang, Lawrence, et al.
Published: (2024)
Towards Scalable Training for Handwritten Mathematical Expression Recognition
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
by: Gupta, Tanmay, et al.
Published: (2026)
by: Gupta, Tanmay, et al.
Published: (2026)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
by: Cho, Janghoon, et al.
Published: (2025)
by: Cho, Janghoon, et al.
Published: (2025)
An Artifact-based Agent Framework for Adaptive and Reproducible Medical Image Processing
by: Zuo, Lianrui, et al.
Published: (2026)
by: Zuo, Lianrui, et al.
Published: (2026)
ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
by: Das, Debasmit, et al.
Published: (2025)
by: Das, Debasmit, et al.
Published: (2025)
Exploration of Reproducible Generated Image Detection
by: Duan, Yihang
Published: (2025)
by: Duan, Yihang
Published: (2025)
WebGuard: Building a Generalizable Guardrail for Web Agents
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
YOLO-Based Pipeline Monitoring in Challenging Visual Environments
by: Dhungana, Pragya, et al.
Published: (2025)
by: Dhungana, Pragya, et al.
Published: (2025)
ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
by: Wang, Kaishen, et al.
Published: (2025)
by: Wang, Kaishen, et al.
Published: (2025)
ViPer: Visual Personalization of Generative Models via Individual Preference Learning
by: Salehi, Sogand, et al.
Published: (2024)
by: Salehi, Sogand, et al.
Published: (2024)
Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models
by: Peng, Ruiying, et al.
Published: (2026)
by: Peng, Ruiying, et al.
Published: (2026)
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
Towards Scalable IoT Deployment for Visual Anomaly Detection via Efficient Compression
by: Stropeni, Arianna, et al.
Published: (2025)
by: Stropeni, Arianna, et al.
Published: (2025)
Similar Items
-
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
by: Bachmann, Roman, et al.
Published: (2024) -
(1D) Ordered Tokens Enable Efficient Test-Time Search
by: Gao, Zhitong, et al.
Published: (2026) -
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
by: Ramachandran, Rahul, et al.
Published: (2025) -
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
by: Atanov, Andrei, et al.
Published: (2026) -
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
by: Bachmann, Roman, et al.
Published: (2025)