WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Miyai, Atsuyuki, Zhao, Zaiying, Egashira, Kazuki, Sato, Atsuki, Sunada, Tatsumi, Onohara, Shota, Yamanishi, Hiromasa, Toyooka, Mashiro, Nishina, Kunato, Maeda, Ryoma, Aizawa, Kiyoharu, Yamasaki, Toshihiko |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
por: Miyai, Atsuyuki, et al.
Publicado: (2026)
por: Miyai, Atsuyuki, et al.
Publicado: (2026)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
por: Miyai, Atsuyuki, et al.
Publicado: (2025)
por: Miyai, Atsuyuki, et al.
Publicado: (2025)
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
por: Miyai, Atsuyuki, et al.
Publicado: (2025)
por: Miyai, Atsuyuki, et al.
Publicado: (2025)
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
por: Baek, Jeonghun, et al.
Publicado: (2026)
por: Baek, Jeonghun, et al.
Publicado: (2026)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
por: Baek, Jeonghun, et al.
Publicado: (2025)
por: Baek, Jeonghun, et al.
Publicado: (2025)
A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
por: Toyooka, Mashiro, et al.
Publicado: (2025)
por: Toyooka, Mashiro, et al.
Publicado: (2025)
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
por: Onohara, Shota, et al.
Publicado: (2024)
por: Onohara, Shota, et al.
Publicado: (2024)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
por: Kawakami, Tatsuki, et al.
Publicado: (2025)
por: Kawakami, Tatsuki, et al.
Publicado: (2025)
GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution Detection
por: Miyai, Atsuyuki, et al.
Publicado: (2023)
por: Miyai, Atsuyuki, et al.
Publicado: (2023)
Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
por: Zhao, Zaiying, et al.
Publicado: (2025)
por: Zhao, Zaiying, et al.
Publicado: (2025)
A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models
por: Noda, Shiho, et al.
Publicado: (2025)
por: Noda, Shiho, et al.
Publicado: (2025)
SVGEditBench: A Benchmark Dataset for Quantitative Assessment of LLM's SVG Editing Capabilities
por: Nishina, Kunato, et al.
Publicado: (2024)
por: Nishina, Kunato, et al.
Publicado: (2024)
SVGEditBench V2: A Benchmark for Instruction-based SVG Editing
por: Nishina, Kunato, et al.
Publicado: (2025)
por: Nishina, Kunato, et al.
Publicado: (2025)
Language-guided Detection and Mitigation of Unknown Dataset Bias
por: Zhao, Zaiying, et al.
Publicado: (2024)
por: Zhao, Zaiying, et al.
Publicado: (2024)
WebGames: Challenging General-Purpose Web-Browsing AI Agents
por: Thomas, George, et al.
Publicado: (2025)
por: Thomas, George, et al.
Publicado: (2025)
A Tool to Facilitate Web-Browsing
por: Kelly, Christopher, et al.
Publicado: (2024)
por: Kelly, Christopher, et al.
Publicado: (2024)
Browsing behavior exposes identities on the Web
por: Oliveira, Marcos, et al.
Publicado: (2023)
por: Oliveira, Marcos, et al.
Publicado: (2023)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
por: Yu, Tao, et al.
Publicado: (2025)
por: Yu, Tao, et al.
Publicado: (2025)
Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
por: Miyai, Atsuyuki, et al.
Publicado: (2024)
por: Miyai, Atsuyuki, et al.
Publicado: (2024)
Beyond Browsing: API-Based Web Agents
por: Song, Yueqi, et al.
Publicado: (2024)
por: Song, Yueqi, et al.
Publicado: (2024)
Go-Browse: Training Web Agents with Structured Exploration
por: Gandhi, Apurva, et al.
Publicado: (2025)
por: Gandhi, Apurva, et al.
Publicado: (2025)
BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
por: Zhou, Peilin, et al.
Publicado: (2025)
por: Zhou, Peilin, et al.
Publicado: (2025)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
por: Lee, Nahyun, et al.
Publicado: (2026)
por: Lee, Nahyun, et al.
Publicado: (2026)
BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
por: Pang, Xianghe, et al.
Publicado: (2025)
por: Pang, Xianghe, et al.
Publicado: (2025)
PixLift: Accelerating Web Browsing via AI Upscaling
por: Atinafu, Yonas, et al.
Publicado: (2025)
por: Atinafu, Yonas, et al.
Publicado: (2025)
Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing
por: Zhang, Shuning, et al.
Publicado: (2025)
por: Zhang, Shuning, et al.
Publicado: (2025)
BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
por: Ou, Litu, et al.
Publicado: (2025)
por: Ou, Litu, et al.
Publicado: (2025)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
por: Miyai, Atsuyuki, et al.
Publicado: (2024)
por: Miyai, Atsuyuki, et al.
Publicado: (2024)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
por: Baek, Jeonghun, et al.
Publicado: (2025)
por: Baek, Jeonghun, et al.
Publicado: (2025)
AiRWeb: Using AR to Extend Web Browsing Beyond Handheld Screens
por: Gao, Mengfei, et al.
Publicado: (2026)
por: Gao, Mengfei, et al.
Publicado: (2026)
Interaction-Driven Browsing: A Human-in-the-Loop Conceptual Framework Informed by Human Web Browsing for Browser-Using Agents
por: Yun, Hyeonggeun, et al.
Publicado: (2025)
por: Yun, Hyeonggeun, et al.
Publicado: (2025)
Prediction of Bone Formation Rate of Artificial Bone With Machine Learning Models Considering the Variation of Experimental Results
por: Yuta Sakai, et al.
Publicado: (2025)
por: Yuta Sakai, et al.
Publicado: (2025)
Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory
por: He, Shiqi, et al.
Publicado: (2025)
por: He, Shiqi, et al.
Publicado: (2025)
Biotic Browser: Applying StreamingLLM as a Persistent Web Browsing Co-Pilot
por: Dunnell, Kevin F., et al.
Publicado: (2024)
por: Dunnell, Kevin F., et al.
Publicado: (2024)
Web-Browsing LLMs Can Access Social Media Profiles and Infer User Demographics
por: Alizadeh, Meysam, et al.
Publicado: (2025)
por: Alizadeh, Meysam, et al.
Publicado: (2025)
Hypertextual Navigation in the SgmlQL Language: Browsing, Querying and Restructuring Web-like Networks
por: Emmanuel Bruno
Publicado: (1999)
por: Emmanuel Bruno
Publicado: (1999)
Read More, Think More: Revisiting Observation Reduction for Web Agents
por: Enomoto, Masafumi, et al.
Publicado: (2026)
por: Enomoto, Masafumi, et al.
Publicado: (2026)
Web Translation Project: WebSpan (the Spanish Web).
por: Ciurczak, Alexis
Publicado: (2000)
por: Ciurczak, Alexis
Publicado: (2000)
Through the Lens of Google CrUX: Dissecting Web Browsing Experience Across Devices and Countries
por: Sengupta, Jayasree, et al.
Publicado: (2023)
por: Sengupta, Jayasree, et al.
Publicado: (2023)
Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Framework
por: Enomoto, Masafumi, et al.
Publicado: (2026)
por: Enomoto, Masafumi, et al.
Publicado: (2026)
Ejemplares similares
-
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
por: Miyai, Atsuyuki, et al.
Publicado: (2026) -
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
por: Miyai, Atsuyuki, et al.
Publicado: (2025) -
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
por: Miyai, Atsuyuki, et al.
Publicado: (2025) -
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
por: Baek, Jeonghun, et al.
Publicado: (2026) -
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
por: Baek, Jeonghun, et al.
Publicado: (2025)