Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Gou, Boyu, Huang, Zanming, Ning, Yuting, Gu, Yu, Lin, Michael, Qi, Weijian, Kopanev, Andrei, Yu, Botao, Gutiérrez, Bernal Jiménez, Shu, Yiheng, Song, Chan Hee, Wu, Jiaman, Chen, Shijie, Moussa, Hanane Nour, Zhang, Tianshu, Xie, Jian, Li, Yifei, Xue, Tianci, Liao, Zeyi, Zhang, Kai, Zheng, Boyuan, Cai, Zhaowei, Rozgic, Viktor, Ziyadi, Morteza, Sun, Huan, Su, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Illusion of Progress? Assessing the Current State of Web Agents
by: Xue, Tianci, et al.
Published: (2025)
by: Xue, Tianci, et al.
Published: (2025)
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
by: Gu, Yu, et al.
Published: (2024)
by: Gu, Yu, et al.
Published: (2024)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
by: Zheng, Boyuan, et al.
Published: (2024)
by: Zheng, Boyuan, et al.
Published: (2024)
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
by: Gou, Boyu, et al.
Published: (2024)
by: Gou, Boyu, et al.
Published: (2024)
JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
by: Bi, Zhenyu, et al.
Published: (2025)
by: Bi, Zhenyu, et al.
Published: (2025)
Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
by: Liao, Zeyi, et al.
Published: (2025)
by: Liao, Zeyi, et al.
Published: (2025)
Secrets of GFlowNets' Learning Behavior: A Theoretical Study
by: Yu, Tianshu
Published: (2025)
by: Yu, Tianshu
Published: (2025)
W2SAT: Learning to generate SAT instances from Weighted Literal Incidence Graphs
by: Wen, Weihuang, et al.
Published: (2023)
by: Wen, Weihuang, et al.
Published: (2023)
D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery
by: Moussa, Hanane Nour, et al.
Published: (2026)
by: Moussa, Hanane Nour, et al.
Published: (2026)
InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
by: Xu, Sirui, et al.
Published: (2026)
by: Xu, Sirui, et al.
Published: (2026)
ChemToolAgent: The Impact of Tools on Language Agents for Chemistry Problem Solving
by: Yu, Botao, et al.
Published: (2024)
by: Yu, Botao, et al.
Published: (2024)
C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs
by: Gao, Rui, et al.
Published: (2026)
by: Gao, Rui, et al.
Published: (2026)
From RAG to Memory: Non-Parametric Continual Learning for Large Language Models
by: Gutiérrez, Bernal Jiménez, et al.
Published: (2025)
by: Gutiérrez, Bernal Jiménez, et al.
Published: (2025)
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
by: Mo, Lingbo, et al.
Published: (2024)
by: Mo, Lingbo, et al.
Published: (2024)
Leveraging Sparse LiDAR for RAFT-Stereo: A Depth Pre-Fill Perspective
by: Yoo, Jinsu, et al.
Published: (2025)
by: Yoo, Jinsu, et al.
Published: (2025)
A Unified Framework for Data-Free One-Step Sampling via Wasserstein Gradient Flows
by: Wang, Chenguang, et al.
Published: (2026)
by: Wang, Chenguang, et al.
Published: (2026)
Learning to Decouple Complex Systems
by: Zhou, Zihan, et al.
Published: (2023)
by: Zhou, Zihan, et al.
Published: (2023)
Certifying Counterfactual Bias in LLMs
by: Chaudhary, Isha, et al.
Published: (2024)
by: Chaudhary, Isha, et al.
Published: (2024)
AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
Realistic threat perception drives intergroup conflict: A causal, dynamic analysis using generative-agent simulations
by: Abdurahman, Suhaib, et al.
Published: (2025)
by: Abdurahman, Suhaib, et al.
Published: (2025)
AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion
by: Li, Hongjie, et al.
Published: (2026)
by: Li, Hongjie, et al.
Published: (2026)
When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
by: Yu, Fangyi
Published: (2025)
by: Yu, Fangyi
Published: (2025)
Regio‐ and Site‐Selective Organic Synthesis With Single‐Atom Catalysts
by: Xin Shang, et al.
Published: (2026)
by: Xin Shang, et al.
Published: (2026)
ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation
by: Li, Hongjie, et al.
Published: (2024)
by: Li, Hongjie, et al.
Published: (2024)
Distributed-downloader
by: Kopanev, Andrei, et al.
Published: (2025)
by: Kopanev, Andrei, et al.
Published: (2025)
Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents
by: Xue, Tianci, et al.
Published: (2026)
by: Xue, Tianci, et al.
Published: (2026)
Data Distribution Bottlenecks in Grounding Language Models to Knowledge Bases
by: Shu, Yiheng, et al.
Published: (2023)
by: Shu, Yiheng, et al.
Published: (2023)
BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
by: Gu, Jianyang, et al.
Published: (2025)
by: Gu, Jianyang, et al.
Published: (2025)
OPTAGENT: Optimizing Multi-Agent LLM Interactions Through Verbal Reinforcement Learning for Enhanced Reasoning
by: Bi, Zhenyu, et al.
Published: (2025)
by: Bi, Zhenyu, et al.
Published: (2025)
RAGPPI: RAG Benchmark for Protein-Protein Interactions in Drug Discovery
by: Jeon, Youngseung, et al.
Published: (2025)
by: Jeon, Youngseung, et al.
Published: (2025)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Lifting Motion to the 3D World via 2D Diffusion
by: Li, Jiaman, et al.
Published: (2024)
by: Li, Jiaman, et al.
Published: (2024)
Codon usage bias of secretory protein in Fusarium oxysporum f. sp. cubense tropical race 4
by: Hui Fang, et al.
Published: (2024)
by: Hui Fang, et al.
Published: (2024)
Progressive Split Mamba: Effective State Space Modelling for Image Restoration
by: Hassanin, Mohammed, et al.
Published: (2026)
by: Hassanin, Mohammed, et al.
Published: (2026)
Achieving Carbon Neutrality for I/O Devices
by: Yu, Botao, et al.
Published: (2024)
by: Yu, Botao, et al.
Published: (2024)
QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks
by: Xie, Jian, et al.
Published: (2026)
by: Xie, Jian, et al.
Published: (2026)
Comprendre la Co-production des Services Publics : Revue des Facteurs Déterminants
by: Hanane Azemzi
Published: (2025)
by: Hanane Azemzi
Published: (2025)
Traduction des parémies marocaines en français : équivalences entre les parémies commençant par « lli » en arabe marocain et par « qui » en français
by: Hanane Hamdane
Published: (2021)
by: Hanane Hamdane
Published: (2021)
Creatividad percibida en problemas insight por alumnado de secundaria: Efectos de algunas variables del problema y del sujeto resolutor
by: Hanane Yousfi
Published: (2024)
by: Hanane Yousfi
Published: (2024)
Model‐Guided Rational Construction of Escherichia coli Synthetic Consortia for Enhanced 2‐Methylbutyric Acid Production
by: Yu Liu, et al.
Published: (2025)
by: Yu Liu, et al.
Published: (2025)
Similar Items
-
An Illusion of Progress? Assessing the Current State of Web Agents
by: Xue, Tianci, et al.
Published: (2025) -
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
by: Gu, Yu, et al.
Published: (2024) -
GPT-4V(ision) is a Generalist Web Agent, if Grounded
by: Zheng, Boyuan, et al.
Published: (2024) -
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
by: Gou, Boyu, et al.
Published: (2024) -
JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
by: Bi, Zhenyu, et al.
Published: (2025)