Automating Benchmark Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dsouza, Amanda, Vishwakarma, Harit, Qi, Zhengyang, Bauer, Justin, Pham, Derek, Walshe, Thomas, Parchami, Armin, Sala, Frederic, Varma, Paroma
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908618675191808
author Dsouza, Amanda
Vishwakarma, Harit
Qi, Zhengyang
Bauer, Justin
Pham, Derek
Walshe, Thomas
Parchami, Armin
Sala, Frederic
Varma, Paroma
author_facet Dsouza, Amanda
Vishwakarma, Harit
Qi, Zhengyang
Bauer, Justin
Pham, Derek
Walshe, Thomas
Parchami, Armin
Sala, Frederic
Varma, Paroma
contents The rapid progress and widespread deployment of LLMs and LLM-powered agents has outpaced our ability to evaluate them. Hand-crafted, static benchmarks are the primary tool for assessing model capabilities, but these quickly become saturated. In contrast, dynamic benchmarks evolve alongside the models they evaluate, but are expensive to create and continuously update. To address these challenges, we develop BeTaL (Benchmark Tuning with an LLM-in-the-loop), a framework that leverages environment design principles to automate the process of dynamic benchmark design. BeTaL works by parameterizing key design choices in base benchmark templates and uses LLMs to reason through the resulting parameter space to obtain target properties (such as difficulty and realism) in a cost-efficient manner. We validate this approach on its ability to create benchmarks with desired difficulty levels. Using BeTaL, we create two new benchmarks and extend a popular agentic benchmark $τ$-bench. Extensive evaluation on these three tasks and multiple target difficulty levels shows that BeTaL produces benchmarks much closer to the desired difficulty, with average deviations ranging from 5.3% to 13.2% -- a 2-4x improvement over the baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automating Benchmark Design
Dsouza, Amanda
Vishwakarma, Harit
Qi, Zhengyang
Bauer, Justin
Pham, Derek
Walshe, Thomas
Parchami, Armin
Sala, Frederic
Varma, Paroma
Software Engineering
Machine Learning
The rapid progress and widespread deployment of LLMs and LLM-powered agents has outpaced our ability to evaluate them. Hand-crafted, static benchmarks are the primary tool for assessing model capabilities, but these quickly become saturated. In contrast, dynamic benchmarks evolve alongside the models they evaluate, but are expensive to create and continuously update. To address these challenges, we develop BeTaL (Benchmark Tuning with an LLM-in-the-loop), a framework that leverages environment design principles to automate the process of dynamic benchmark design. BeTaL works by parameterizing key design choices in base benchmark templates and uses LLMs to reason through the resulting parameter space to obtain target properties (such as difficulty and realism) in a cost-efficient manner. We validate this approach on its ability to create benchmarks with desired difficulty levels. Using BeTaL, we create two new benchmarks and extend a popular agentic benchmark $τ$-bench. Extensive evaluation on these three tasks and multiple target difficulty levels shows that BeTaL produces benchmarks much closer to the desired difficulty, with average deviations ranging from 5.3% to 13.2% -- a 2-4x improvement over the baselines.
title Automating Benchmark Design
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2510.25039