LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: White, Colin, Dooley, Samuel, Roberts, Manley, Pal, Arka, Feuer, Ben, Jain, Siddhartha, Shwartz-Ziv, Ravid, Jain, Neel, Saifullah, Khalid, Dey, Sreemanti, Shubh-Agrawal, Sandha, Sandeep Singh, Naidu, Siddartha, Hegde, Chinmay, LeCun, Yann, Goldstein, Tom, Neiswanger, Willie, Goldblum, Micah
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915250078482432
author White, Colin
Dooley, Samuel
Roberts, Manley
Pal, Arka
Feuer, Ben
Jain, Siddhartha
Shwartz-Ziv, Ravid
Jain, Neel
Saifullah, Khalid
Dey, Sreemanti
Shubh-Agrawal
Sandha, Sandeep Singh
Naidu, Siddartha
Hegde, Chinmay
LeCun, Yann
Goldstein, Tom
Neiswanger, Willie
Goldblum, Micah
author_facet White, Colin
Dooley, Samuel
Roberts, Manley
Pal, Arka
Feuer, Ben
Jain, Siddhartha
Shwartz-Ziv, Ravid
Jain, Neel
Saifullah, Khalid
Dey, Sreemanti
Shubh-Agrawal
Sandha, Sandeep Singh
Naidu, Siddartha
Hegde, Chinmay
LeCun, Yann
Goldstein, Tom
Neiswanger, Willie
Goldblum, Micah
contents Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19314
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LiveBench: A Challenging, Contamination-Limited LLM Benchmark
White, Colin
Dooley, Samuel
Roberts, Manley
Pal, Arka
Feuer, Ben
Jain, Siddhartha
Shwartz-Ziv, Ravid
Jain, Neel
Saifullah, Khalid
Dey, Sreemanti
Shubh-Agrawal
Sandha, Sandeep Singh
Naidu, Siddartha
Hegde, Chinmay
LeCun, Yann
Goldstein, Tom
Neiswanger, Willie
Goldblum, Micah
Computation and Language
Artificial Intelligence
Machine Learning
Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models.
title LiveBench: A Challenging, Contamination-Limited LLM Benchmark
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.19314