WebSight: A Vision-First Architecture for Robust Web Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Bhathal, Tanvir, Gupta, Asanshay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
WebGuard: Building a Generalizable Guardrail for Web Agents
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
by: Jang, Lawrence, et al.
Published: (2024)
by: Jang, Lawrence, et al.
Published: (2024)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
by: Zhang, Ziyun, et al.
Published: (2026)
by: Zhang, Ziyun, et al.
Published: (2026)
PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
WebInject: Prompt Injection Attack to Web Agents
by: Wang, Xilong, et al.
Published: (2025)
by: Wang, Xilong, et al.
Published: (2025)
WALT: Web Agents that Learn Tools
by: Prabhu, Viraj, et al.
Published: (2025)
by: Prabhu, Viraj, et al.
Published: (2025)
WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark
by: Yuan, Peng, et al.
Published: (2026)
by: Yuan, Peng, et al.
Published: (2026)
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
by: Fan, Sicheng, et al.
Published: (2026)
by: Fan, Sicheng, et al.
Published: (2026)
Towards Safe Synthetic Image Generation On the Web: A Multimodal Robust NSFW Defense and Million Scale Dataset
by: Muneer, Muhammad Shahid, et al.
Published: (2025)
by: Muneer, Muhammad Shahid, et al.
Published: (2025)
Web World Models
by: Feng, Jichen, et al.
Published: (2025)
by: Feng, Jichen, et al.
Published: (2025)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
by: Jia, Yiming, et al.
Published: (2025)
by: Jia, Yiming, et al.
Published: (2025)
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
InkSight: Offline-to-Online Handwriting Conversion by Teaching Vision-Language Models to Read and Write
by: Mitrevski, Blagoj, et al.
Published: (2024)
by: Mitrevski, Blagoj, et al.
Published: (2024)
WebAccessVL: Violation-Aware VLM for Web Accessibility
by: Zheng, Amber Yijia, et al.
Published: (2025)
by: Zheng, Amber Yijia, et al.
Published: (2025)
Vision Learners Meet Web Image-Text Pairs
by: Zhao, Bingchen, et al.
Published: (2023)
by: Zhao, Bingchen, et al.
Published: (2023)
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
by: Xie, Rui, et al.
Published: (2026)
by: Xie, Rui, et al.
Published: (2026)
TimeWarp: Evaluating Web Agents by Revisiting the Past
by: Ishmam, Md Farhan, et al.
Published: (2026)
by: Ishmam, Md Farhan, et al.
Published: (2026)
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
by: Liu, Zhihong, et al.
Published: (2026)
by: Liu, Zhihong, et al.
Published: (2026)
Line of Sight: On Linear Representations in VLLMs
by: Rajaram, Achyuta, et al.
Published: (2025)
by: Rajaram, Achyuta, et al.
Published: (2025)
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
by: Hong, Yining, et al.
Published: (2025)
by: Hong, Yining, et al.
Published: (2025)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026)
by: Guo, JunJia, et al.
Published: (2026)
Self-supervised visual learning for analyzing firearms trafficking activities on the Web
by: Konstantakos, Sotirios, et al.
Published: (2023)
by: Konstantakos, Sotirios, et al.
Published: (2023)
A Survey on Mamba Architecture for Vision Applications
by: Ibrahim, Fady, et al.
Published: (2025)
by: Ibrahim, Fady, et al.
Published: (2025)
Dual-View Visual Contextualization for Web Navigation
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
by: Xu, Kevin, et al.
Published: (2024)
by: Xu, Kevin, et al.
Published: (2024)
EEO-TFV: Escape-Explore Optimizer for Web-Scale Time-Series Forecasting and Vision Analysis
by: Wang, Hua, et al.
Published: (2026)
by: Wang, Hua, et al.
Published: (2026)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
by: Zheng, Boyuan, et al.
Published: (2024)
by: Zheng, Boyuan, et al.
Published: (2024)
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
by: Liu, Chengwen, et al.
Published: (2026)
by: Liu, Chengwen, et al.
Published: (2026)
WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale Benchmark
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
by: Ye, Suyu, et al.
Published: (2025)
by: Ye, Suyu, et al.
Published: (2025)
OOSTraj: Out-of-Sight Trajectory Prediction With Vision-Positioning Denoising
by: Zhang, Haichao, et al.
Published: (2024)
by: Zhang, Haichao, et al.
Published: (2024)
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding
by: Bai, Yatong, et al.
Published: (2024)
by: Bai, Yatong, et al.
Published: (2024)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Similar Items
-
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
by: Laurençon, Hugo, et al.
Published: (2024) -
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026) -
WebGuard: Building a Generalizable Guardrail for Web Agents
by: Zheng, Boyuan, et al.
Published: (2025) -
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026) -
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
by: Jang, Lawrence, et al.
Published: (2024)