From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Tianle, Chiang, Wei-Lin, Frick, Evan, Dunlap, Lisa, Wu, Tianhao, Zhu, Banghua, Gonzalez, Joseph E., Stoica, Ion
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913545452519424
author Li, Tianle
Chiang, Wei-Lin
Frick, Evan
Dunlap, Lisa
Wu, Tianhao
Zhu, Banghua
Gonzalez, Joseph E.
Stoica, Ion
author_facet Li, Tianle
Chiang, Wei-Lin
Frick, Evan
Dunlap, Lisa
Wu, Tianhao
Zhu, Banghua
Gonzalez, Joseph E.
Stoica, Ion
contents The rapid evolution of Large Language Models (LLMs) has outpaced the development of model evaluation, highlighting the need for continuous curation of new, challenging benchmarks. However, manual curation of high-quality, human-aligned benchmarks is expensive and time-consuming. To address this, we introduce BenchBuilder, an automated pipeline that leverages LLMs to curate high-quality, open-ended prompts from large, crowd-sourced datasets, enabling continuous benchmark updates without human in the loop. We apply BenchBuilder to datasets such as Chatbot Arena and WildChat-1M, extracting challenging prompts and utilizing LLM-as-a-Judge for automatic model evaluation. To validate benchmark quality, we propose new metrics to measure a benchmark's alignment with human preferences and ability to separate models. We release Arena-Hard-Auto, a benchmark consisting 500 challenging prompts curated by BenchBuilder. Arena-Hard-Auto provides 3x higher separation of model performances compared to MT-Bench and achieves 98.6% correlation with human preference rankings, all at a cost of $20. Our work sets a new framework for the scalable curation of automated benchmarks from extensive data.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Li, Tianle
Chiang, Wei-Lin
Frick, Evan
Dunlap, Lisa
Wu, Tianhao
Zhu, Banghua
Gonzalez, Joseph E.
Stoica, Ion
Machine Learning
Artificial Intelligence
Computation and Language
The rapid evolution of Large Language Models (LLMs) has outpaced the development of model evaluation, highlighting the need for continuous curation of new, challenging benchmarks. However, manual curation of high-quality, human-aligned benchmarks is expensive and time-consuming. To address this, we introduce BenchBuilder, an automated pipeline that leverages LLMs to curate high-quality, open-ended prompts from large, crowd-sourced datasets, enabling continuous benchmark updates without human in the loop. We apply BenchBuilder to datasets such as Chatbot Arena and WildChat-1M, extracting challenging prompts and utilizing LLM-as-a-Judge for automatic model evaluation. To validate benchmark quality, we propose new metrics to measure a benchmark's alignment with human preferences and ability to separate models. We release Arena-Hard-Auto, a benchmark consisting 500 challenging prompts curated by BenchBuilder. Arena-Hard-Auto provides 3x higher separation of model performances compared to MT-Bench and achieves 98.6% correlation with human preference rankings, all at a cost of $20. Our work sets a new framework for the scalable curation of automated benchmarks from extensive data.
title From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2406.11939