AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
Fuente:
arXiv
Saved in:
| Main Authors: | Lù, Xing Han, Kazemnejad, Amirhossein, Meade, Nicholas, Patel, Arkil, Shin, Dongchan, Zambrano, Alejandra, Stańczak, Karolina, Shaw, Peter, Pal, Christopher J., Reddy, Siva |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
Investigating Adversarial Trigger Transfer in Large Language Models
by: Meade, Nicholas, et al.
Published: (2024)
by: Meade, Nicholas, et al.
Published: (2024)
How to Get Your LLM to Generate Challenging Problems for Evaluation
by: Patel, Arkil, et al.
Published: (2025)
by: Patel, Arkil, et al.
Published: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025)
by: Marjanović, Sara Vera, et al.
Published: (2025)
Evaluating In-Context Learning of Libraries for Code Generation
by: Patel, Arkil, et al.
Published: (2023)
by: Patel, Arkil, et al.
Published: (2023)
Structured Distillation of Web Agent Capabilities Enables Generalization
by: Lù, Xing Han, et al.
Published: (2026)
by: Lù, Xing Han, et al.
Published: (2026)
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026)
by: Patel, Arkil, et al.
Published: (2026)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
by: Adlakha, Vaibhav, et al.
Published: (2023)
by: Adlakha, Vaibhav, et al.
Published: (2023)
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
by: BehnamGhader, Parishad, et al.
Published: (2025)
by: BehnamGhader, Parishad, et al.
Published: (2025)
A Multilingual Perspective on Probing Gender Bias
by: Stańczak, Karolina
Published: (2024)
by: Stańczak, Karolina
Published: (2024)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
by: Aghajohari, Milad, et al.
Published: (2025)
by: Aghajohari, Milad, et al.
Published: (2025)
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
by: Levy, Ido, et al.
Published: (2024)
by: Levy, Ido, et al.
Published: (2024)
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
by: Ding, Xuwei, et al.
Published: (2026)
by: Ding, Xuwei, et al.
Published: (2026)
CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
by: Lin, Haojia, et al.
Published: (2025)
by: Lin, Haojia, et al.
Published: (2025)
VinePPO: Refining Credit Assignment in RL Training of LLMs
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection
by: Zadeh, Fatemeh Pesaran, et al.
Published: (2026)
by: Zadeh, Fatemeh Pesaran, et al.
Published: (2026)
Deep Research Bench: Evaluating AI Web Research Agents
by: FutureSearch, et al.
Published: (2025)
by: FutureSearch, et al.
Published: (2025)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025)
by: Bhatia, Mehar, et al.
Published: (2025)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
by: Lù, Xing Han, et al.
Published: (2024)
by: Lù, Xing Han, et al.
Published: (2024)
WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
by: Qian, Yaoyao, et al.
Published: (2025)
by: Qian, Yaoyao, et al.
Published: (2025)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
by: Kong, Fanheng, et al.
Published: (2026)
by: Kong, Fanheng, et al.
Published: (2026)
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
by: Zhong, Shanshan, et al.
Published: (2026)
by: Zhong, Shanshan, et al.
Published: (2026)
Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents
by: Zambrano, Alejandra, et al.
Published: (2026)
by: Zambrano, Alejandra, et al.
Published: (2026)
FIRE-Bench: Evaluating Agents on the Rediscovery of Scientific Insights
by: Wang, Zhen, et al.
Published: (2026)
by: Wang, Zhen, et al.
Published: (2026)
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
by: Logeswaran, Lajanugen, et al.
Published: (2026)
by: Logeswaran, Lajanugen, et al.
Published: (2026)
BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
by: Fa, Dionizije, et al.
Published: (2026)
by: Fa, Dionizije, et al.
Published: (2026)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
by: Zheng, Junhao, et al.
Published: (2025)
by: Zheng, Junhao, et al.
Published: (2025)
The Promise of RL for Autoregressive Image Editing
by: Ahmadi, Saba, et al.
Published: (2025)
by: Ahmadi, Saba, et al.
Published: (2025)
SocialBench: Sociality Evaluation of Role-Playing Conversational Agents
by: Chen, Hongzhan, et al.
Published: (2024)
by: Chen, Hongzhan, et al.
Published: (2024)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
by: Saxena, Siddhant, et al.
Published: (2026)
by: Saxena, Siddhant, et al.
Published: (2026)
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games
by: Zhang, Wenyu, et al.
Published: (2026)
by: Zhang, Wenyu, et al.
Published: (2026)
Evaluating and Improving Graph-based Explanation Methods for Multi-Agent Coordination
by: Kailas, Siva, et al.
Published: (2025)
by: Kailas, Siva, et al.
Published: (2025)
CocoaBench: Evaluating Unified Digital Agents in the Wild
by: CocoaBench Team, et al.
Published: (2026)
by: CocoaBench Team, et al.
Published: (2026)
RewardBench 2: Advancing Reward Model Evaluation
by: Malik, Saumya, et al.
Published: (2025)
by: Malik, Saumya, et al.
Published: (2025)
RewardBench: Evaluating Reward Models for Language Modeling
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
Similar Items
-
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025) -
Investigating Adversarial Trigger Transfer in Large Language Models
by: Meade, Nicholas, et al.
Published: (2024) -
How to Get Your LLM to Generate Challenging Problems for Evaluation
by: Patel, Arkil, et al.
Published: (2025) -
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025) -
Evaluating In-Context Learning of Libraries for Code Generation
by: Patel, Arkil, et al.
Published: (2023)