WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Peeters, Ralph, Steiner, Aaron, Schwarz, Luca, Caspary, Julian Yuya, Bizer, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report)
by: Steiner, Aaron, et al.
Published: (2025)
by: Steiner, Aaron, et al.
Published: (2025)
Entity Matching using Large Language Models
by: Peeters, Ralph, et al.
Published: (2023)
by: Peeters, Ralph, et al.
Published: (2023)
Fine-tuning Large Language Models for Entity Matching
by: Steiner, Aaron, et al.
Published: (2024)
by: Steiner, Aaron, et al.
Published: (2024)
Automatic End-to-End Data Integration using Large Language Models
by: Steiner, Aaron, et al.
Published: (2026)
by: Steiner, Aaron, et al.
Published: (2026)
DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
by: Zhang, Enze, et al.
Published: (2025)
by: Zhang, Enze, et al.
Published: (2025)
Evaluating Knowledge Generation and Self-Refinement Strategies for LLM-based Column Type Annotation
by: Korini, Keti, et al.
Published: (2025)
by: Korini, Keti, et al.
Published: (2025)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Self-Refinement Strategies for LLM-based Product Attribute Value Extraction
by: Brinkmann, Alexander, et al.
Published: (2025)
by: Brinkmann, Alexander, et al.
Published: (2025)
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments
by: Liu, Zefang, et al.
Published: (2025)
by: Liu, Zefang, et al.
Published: (2025)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
by: Zhang, Xianren, et al.
Published: (2025)
by: Zhang, Xianren, et al.
Published: (2025)
WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code
by: Lin, Zhiyu, et al.
Published: (2025)
by: Lin, Zhiyu, et al.
Published: (2025)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
by: Kim, Serin, et al.
Published: (2026)
by: Kim, Serin, et al.
Published: (2026)
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
by: Miyai, Atsuyuki, et al.
Published: (2025)
by: Miyai, Atsuyuki, et al.
Published: (2025)
LiveWeb-IE: A Benchmark For Online Web Information Extraction
by: Yang, Seungbin, et al.
Published: (2026)
by: Yang, Seungbin, et al.
Published: (2026)
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
by: Wang, Jiangyuan, et al.
Published: (2025)
by: Wang, Jiangyuan, et al.
Published: (2025)
WebDS: An End-to-End Benchmark for Web-based Data Science
by: Hsu, Ethan, et al.
Published: (2025)
by: Hsu, Ethan, et al.
Published: (2025)
WCXB: A Multi-Type Web Content Extraction Benchmark
by: Foley, Murrough
Published: (2026)
by: Foley, Murrough
Published: (2026)
WebCanvas: Benchmarking Web Agents in Online Environments
by: Pan, Yichen, et al.
Published: (2024)
by: Pan, Yichen, et al.
Published: (2024)
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
by: Chae, Hyungjoo, et al.
Published: (2025)
by: Chae, Hyungjoo, et al.
Published: (2025)
WebWalker: Benchmarking LLMs in Web Traversal
by: Wu, Jialong, et al.
Published: (2025)
by: Wu, Jialong, et al.
Published: (2025)
Evaluating Cultural and Social Awareness of LLM Web Agents
by: Qiu, Haoyi, et al.
Published: (2024)
by: Qiu, Haoyi, et al.
Published: (2024)
Using LLMs for the Extraction and Normalization of Product Attribute Values
by: Brinkmann, Alexander, et al.
Published: (2024)
by: Brinkmann, Alexander, et al.
Published: (2024)
ExtractGPT: Exploring the Potential of Large Language Models for Product Attribute Value Extraction
by: Brinkmann, Alexander, et al.
Published: (2023)
by: Brinkmann, Alexander, et al.
Published: (2023)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
by: Wang, Yumeng, et al.
Published: (2025)
by: Wang, Yumeng, et al.
Published: (2025)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025)
by: Wadhwa, Manya, et al.
Published: (2025)
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
by: Wei, Zhepei, et al.
Published: (2025)
by: Wei, Zhepei, et al.
Published: (2025)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
by: Jang, Lawrence Keunho, et al.
Published: (2026)
by: Jang, Lawrence Keunho, et al.
Published: (2026)
WebXSkill: Skill Learning for Autonomous Web Agents
by: Wang, Zhaoyang, et al.
Published: (2026)
by: Wang, Zhaoyang, et al.
Published: (2026)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
GuideWeb: A Benchmark for Automatic In-App Guide Generation on Real-World Web UIs
by: Gan, Chengguang, et al.
Published: (2026)
by: Gan, Chengguang, et al.
Published: (2026)
AutoWebGLM: A Large Language Model-based Web Navigating Agent
by: Lai, Hanyu, et al.
Published: (2024)
by: Lai, Hanyu, et al.
Published: (2024)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
by: Fang, Tianqing, et al.
Published: (2025)
by: Fang, Tianqing, et al.
Published: (2025)
WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning
by: Yu, Xinmiao, et al.
Published: (2026)
by: Yu, Xinmiao, et al.
Published: (2026)
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
by: Chae, Hyungjoo, et al.
Published: (2024)
by: Chae, Hyungjoo, et al.
Published: (2024)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
by: Yu, Tao, et al.
Published: (2025)
by: Yu, Tao, et al.
Published: (2025)
WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
by: Zhang, Zhisong, et al.
Published: (2025)
by: Zhang, Zhisong, et al.
Published: (2025)
Similar Items
-
MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report)
by: Steiner, Aaron, et al.
Published: (2025) -
Entity Matching using Large Language Models
by: Peeters, Ralph, et al.
Published: (2023) -
Fine-tuning Large Language Models for Entity Matching
by: Steiner, Aaron, et al.
Published: (2024) -
Automatic End-to-End Data Integration using Large Language Models
by: Steiner, Aaron, et al.
Published: (2026) -
DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
by: Zhang, Enze, et al.
Published: (2025)