Towards a Realistic Long-Term Benchmark for Open-Web Research Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mühlbacher, Peter, Bosse, Nikos I., Phillips, Lawrence
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909325142786048
author Mühlbacher, Peter
Bosse, Nikos I.
Phillips, Lawrence
author_facet Mühlbacher, Peter
Bosse, Nikos I.
Phillips, Lawrence
contents We present initial results of a forthcoming benchmark for evaluating LLM agents on white-collar tasks of economic value. We evaluate agents on real-world "messy" open-web research tasks of the type that are routine in finance and consulting. In doing so, we lay the groundwork for an LLM agent evaluation suite where good performance directly corresponds to a large economic and societal impact. We built and tested several agent architectures with o1-preview, GPT-4o, Claude-3.5 Sonnet, Llama 3.1 (405b), and GPT-4o-mini. On average, LLM agents powered by Claude-3.5 Sonnet and o1-preview substantially outperformed agents using GPT-4o, with agents based on Llama 3.1 (405b) and GPT-4o-mini lagging noticeably behind. Across LLMs, a ReAct architecture with the ability to delegate subtasks to subagents performed best. In addition to quantitative evaluations, we qualitatively assessed the performance of the LLM agents by inspecting their traces and reflecting on their observations. Our evaluation represents the first in-depth assessment of agents' abilities to conduct challenging, economically valuable analyst-style research on the real open web.
format Preprint
id arxiv_https___arxiv_org_abs_2409_14913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards a Realistic Long-Term Benchmark for Open-Web Research Agents
Mühlbacher, Peter
Bosse, Nikos I.
Phillips, Lawrence
Computation and Language
Information Retrieval
Machine Learning
We present initial results of a forthcoming benchmark for evaluating LLM agents on white-collar tasks of economic value. We evaluate agents on real-world "messy" open-web research tasks of the type that are routine in finance and consulting. In doing so, we lay the groundwork for an LLM agent evaluation suite where good performance directly corresponds to a large economic and societal impact. We built and tested several agent architectures with o1-preview, GPT-4o, Claude-3.5 Sonnet, Llama 3.1 (405b), and GPT-4o-mini. On average, LLM agents powered by Claude-3.5 Sonnet and o1-preview substantially outperformed agents using GPT-4o, with agents based on Llama 3.1 (405b) and GPT-4o-mini lagging noticeably behind. Across LLMs, a ReAct architecture with the ability to delegate subtasks to subagents performed best. In addition to quantitative evaluations, we qualitatively assessed the performance of the LLM agents by inspecting their traces and reflecting on their observations. Our evaluation represents the first in-depth assessment of agents' abilities to conduct challenging, economically valuable analyst-style research on the real open web.
title Towards a Realistic Long-Term Benchmark for Open-Web Research Agents
topic Computation and Language
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2409.14913