LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Peiwen, Feng, Shaoxiong, Li, Yiwei, Wang, Xinglin, Zhang, Yueqi, Shi, Jiayi, Tan, Chuyi, Pan, Boyuan, Hu, Yao, Li, Kan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929696782942208
author Yuan, Peiwen
Feng, Shaoxiong
Li, Yiwei
Wang, Xinglin
Zhang, Yueqi
Shi, Jiayi
Tan, Chuyi
Pan, Boyuan
Hu, Yao
Li, Kan
author_facet Yuan, Peiwen
Feng, Shaoxiong
Li, Yiwei
Wang, Xinglin
Zhang, Yueqi
Shi, Jiayi
Tan, Chuyi
Pan, Boyuan
Hu, Yao
Li, Kan
contents The rapid advancement of large language models (LLMs) has led to a surge in both model supply and application demands. To facilitate effective matching between them, reliable, generic and efficient benchmark generators are widely needed. However, human annotators are constrained by inefficiency, and current LLM benchmark generators not only lack generalizability but also struggle with limited reliability, as they lack a comprehensive evaluation framework for validation and optimization. To fill this gap, we first propose an automated and unbiased evaluation framework, structured around four dimensions and ten criteria. Under this framework, we carefully analyze the advantages and weaknesses of directly prompting LLMs as generic benchmark generators. To enhance the reliability, we introduce a series of methods to address the identified weaknesses and integrate them as BenchMaker. Experiments across multiple LLMs and tasks confirm that BenchMaker achieves superior or comparable performance to human-annotated benchmarks on all metrics, highlighting its generalizability and reliability. More importantly, it delivers highly consistent evaluation results across 12 LLMs (0.967 Pearson correlation against MMLU-Pro), while taking only $0.005 and 0.38 minutes per sample.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01683
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
Yuan, Peiwen
Feng, Shaoxiong
Li, Yiwei
Wang, Xinglin
Zhang, Yueqi
Shi, Jiayi
Tan, Chuyi
Pan, Boyuan
Hu, Yao
Li, Kan
Computation and Language
Artificial Intelligence
The rapid advancement of large language models (LLMs) has led to a surge in both model supply and application demands. To facilitate effective matching between them, reliable, generic and efficient benchmark generators are widely needed. However, human annotators are constrained by inefficiency, and current LLM benchmark generators not only lack generalizability but also struggle with limited reliability, as they lack a comprehensive evaluation framework for validation and optimization. To fill this gap, we first propose an automated and unbiased evaluation framework, structured around four dimensions and ten criteria. Under this framework, we carefully analyze the advantages and weaknesses of directly prompting LLMs as generic benchmark generators. To enhance the reliability, we introduce a series of methods to address the identified weaknesses and integrate them as BenchMaker. Experiments across multiple LLMs and tasks confirm that BenchMaker achieves superior or comparable performance to human-annotated benchmarks on all metrics, highlighting its generalizability and reliability. More importantly, it delivers highly consistent evaluation results across 12 LLMs (0.967 Pearson correlation against MMLU-Pro), while taking only $0.005 and 0.38 minutes per sample.
title LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.01683