GuideWeb: A Benchmark for Automatic In-App Guide Generation on Real-World Web UIs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gan, Chengguang, Tsujii, Yoshihiro, Liang, Yunhao, Mori, Tatsunori, Ni, Shiwen, Itoh, Hiroki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
Empirical Study of Mutual Reinforcement Effect and Application in Few-shot Text Classification Tasks via Prompt
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
Application of LLM Agents in Recruitment: A Novel Framework for Resume Screening
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
Demonstrating Mutual Reinforcement Effect through Information Flow
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
A Multilingual Dataset and Empirical Validation for the Mutual Reinforcement Effect in Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
von: Ye, Suyu, et al.
Veröffentlicht: (2025)
von: Ye, Suyu, et al.
Veröffentlicht: (2025)
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
von: Gur, Izzeddin, et al.
Veröffentlicht: (2023)
von: Gur, Izzeddin, et al.
Veröffentlicht: (2023)
NaviQAte: Functionality-Guided Web Application Navigation
von: Shahbandeh, Mobina, et al.
Veröffentlicht: (2024)
von: Shahbandeh, Mobina, et al.
Veröffentlicht: (2024)
WebWalker: Benchmarking LLMs in Web Traversal
von: Wu, Jialong, et al.
Veröffentlicht: (2025)
von: Wu, Jialong, et al.
Veröffentlicht: (2025)
Automatic Generation of Web Censorship Probe Lists
von: Tang, Jenny, et al.
Veröffentlicht: (2024)
von: Tang, Jenny, et al.
Veröffentlicht: (2024)
MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
LiveWeb-IE: A Benchmark For Online Web Information Extraction
von: Yang, Seungbin, et al.
Veröffentlicht: (2026)
von: Yang, Seungbin, et al.
Veröffentlicht: (2026)
WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
von: Peeters, Ralph, et al.
Veröffentlicht: (2025)
von: Peeters, Ralph, et al.
Veröffentlicht: (2025)
OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
von: He, Hongliang, et al.
Veröffentlicht: (2024)
von: He, Hongliang, et al.
Veröffentlicht: (2024)
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
von: Xu, Yiheng, et al.
Veröffentlicht: (2024)
von: Xu, Yiheng, et al.
Veröffentlicht: (2024)
WRAP++: Web discoveRy Amplified Pretraining
von: Zhou, Jiang, et al.
Veröffentlicht: (2026)
von: Zhou, Jiang, et al.
Veröffentlicht: (2026)
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2024)
von: Chae, Hyungjoo, et al.
Veröffentlicht: (2024)
Harnessing Webpage UIs for Text-Rich Visual Understanding
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
von: Ou, Litu, et al.
Veröffentlicht: (2025)
von: Ou, Litu, et al.
Veröffentlicht: (2025)
Web World Models
von: Feng, Jichen, et al.
Veröffentlicht: (2025)
von: Feng, Jichen, et al.
Veröffentlicht: (2025)
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
von: Wang, Yumeng, et al.
Veröffentlicht: (2025)
von: Wang, Yumeng, et al.
Veröffentlicht: (2025)
WebDS: An End-to-End Benchmark for Web-based Data Science
von: Hsu, Ethan, et al.
Veröffentlicht: (2025)
von: Hsu, Ethan, et al.
Veröffentlicht: (2025)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
von: Lù, Xing Han, et al.
Veröffentlicht: (2024)
von: Lù, Xing Han, et al.
Veröffentlicht: (2024)
AutoScraper: A Progressive Understanding Web Agent for Web Scraper Generation
von: Huang, Wenhao, et al.
Veröffentlicht: (2024)
von: Huang, Wenhao, et al.
Veröffentlicht: (2024)
Educational-Psychological Dialogue Robot Based on Multi-Agent Collaboration
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code
von: Lin, Zhiyu, et al.
Veröffentlicht: (2025)
von: Lin, Zhiyu, et al.
Veröffentlicht: (2025)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
von: Wang, Peng, et al.
Veröffentlicht: (2025)
von: Wang, Peng, et al.
Veröffentlicht: (2025)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
von: Kim, Serin, et al.
Veröffentlicht: (2026)
von: Kim, Serin, et al.
Veröffentlicht: (2026)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
WeKnow-RAG: An Adaptive Approach for Retrieval-Augmented Generation Integrating Web Search and Knowledge Graphs
von: Xie, Weijian, et al.
Veröffentlicht: (2024)
von: Xie, Weijian, et al.
Veröffentlicht: (2024)
A Case Study of Web App Coding with OpenAI Reasoning Models
von: Cui, Yi
Veröffentlicht: (2024)
von: Cui, Yi
Veröffentlicht: (2024)
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
WCXB: A Multi-Type Web Content Extraction Benchmark
von: Foley, Murrough
Veröffentlicht: (2026)
von: Foley, Murrough
Veröffentlicht: (2026)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2026)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2026)
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
von: Dubois, Yann, et al.
Veröffentlicht: (2024)
von: Dubois, Yann, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2025) -
Empirical Study of Mutual Reinforcement Effect and Application in Few-shot Text Classification Tasks via Prompt
von: Gan, Chengguang, et al.
Veröffentlicht: (2024) -
Application of LLM Agents in Recruitment: A Novel Framework for Resume Screening
von: Gan, Chengguang, et al.
Veröffentlicht: (2024) -
Demonstrating Mutual Reinforcement Effect through Information Flow
von: Gan, Chengguang, et al.
Veröffentlicht: (2024) -
A Multilingual Dataset and Empirical Validation for the Mutual Reinforcement Effect in Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)