The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Seungone, Suk, Juyoung, Cho, Ji Yong, Longpre, Shayne, Kim, Chaeeun, Yoon, Dongkeun, Son, Guijin, Cho, Yejin, Shafayat, Sheikh, Baek, Jinheon, Park, Sue Hyun, Hwang, Hyeonbin, Jo, Jinkyung, Cho, Hyowon, Shin, Haebin, Lee, Seongyun, Oh, Hanseok, Lee, Noah, Ho, Namgyu, Joo, Se June, Ko, Miyoung, Lee, Yoonjoo, Chae, Hyungjoo, Shin, Jamin, Jang, Joel, Ye, Seonghyeon, Lin, Bill Yuchen, Welleck, Sean, Neubig, Graham, Lee, Moontae, Lee, Kyungjae, Seo, Minjoon
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910890589159424
author Kim, Seungone
Suk, Juyoung
Cho, Ji Yong
Longpre, Shayne
Kim, Chaeeun
Yoon, Dongkeun
Son, Guijin
Cho, Yejin
Shafayat, Sheikh
Baek, Jinheon
Park, Sue Hyun
Hwang, Hyeonbin
Jo, Jinkyung
Cho, Hyowon
Shin, Haebin
Lee, Seongyun
Oh, Hanseok
Lee, Noah
Ho, Namgyu
Joo, Se June
Ko, Miyoung
Lee, Yoonjoo
Chae, Hyungjoo
Shin, Jamin
Jang, Joel
Ye, Seonghyeon
Lin, Bill Yuchen
Welleck, Sean
Neubig, Graham
Lee, Moontae
Lee, Kyungjae
Seo, Minjoon
author_facet Kim, Seungone
Suk, Juyoung
Cho, Ji Yong
Longpre, Shayne
Kim, Chaeeun
Yoon, Dongkeun
Son, Guijin
Cho, Yejin
Shafayat, Sheikh
Baek, Jinheon
Park, Sue Hyun
Hwang, Hyeonbin
Jo, Jinkyung
Cho, Hyowon
Shin, Haebin
Lee, Seongyun
Oh, Hanseok
Lee, Noah
Ho, Namgyu
Joo, Se June
Ko, Miyoung
Lee, Yoonjoo
Chae, Hyungjoo
Shin, Jamin
Jang, Joel
Ye, Seonghyeon
Lin, Bill Yuchen
Welleck, Sean
Neubig, Graham
Lee, Moontae
Lee, Kyungjae
Seo, Minjoon
contents As language models (LMs) become capable of handling a wide range of tasks, their evaluation is becoming as challenging as their development. Most generation benchmarks currently assess LMs using abstract evaluation criteria like helpfulness and harmlessness, which often lack the flexibility and granularity of human assessment. Additionally, these benchmarks tend to focus disproportionately on specific capabilities such as instruction following, leading to coverage bias. To overcome these limitations, we introduce the BiGGen Bench, a principled generation benchmark designed to thoroughly evaluate nine distinct capabilities of LMs across 77 diverse tasks. A key feature of the BiGGen Bench is its use of instance-specific evaluation criteria, closely mirroring the nuanced discernment of human evaluation. We apply this benchmark to assess 103 frontier LMs using five evaluator LMs. Our code, data, and evaluation results are all publicly available at https://github.com/prometheus-eval/prometheus-eval/tree/main/BiGGen-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05761
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
Kim, Seungone
Suk, Juyoung
Cho, Ji Yong
Longpre, Shayne
Kim, Chaeeun
Yoon, Dongkeun
Son, Guijin
Cho, Yejin
Shafayat, Sheikh
Baek, Jinheon
Park, Sue Hyun
Hwang, Hyeonbin
Jo, Jinkyung
Cho, Hyowon
Shin, Haebin
Lee, Seongyun
Oh, Hanseok
Lee, Noah
Ho, Namgyu
Joo, Se June
Ko, Miyoung
Lee, Yoonjoo
Chae, Hyungjoo
Shin, Jamin
Jang, Joel
Ye, Seonghyeon
Lin, Bill Yuchen
Welleck, Sean
Neubig, Graham
Lee, Moontae
Lee, Kyungjae
Seo, Minjoon
Computation and Language
As language models (LMs) become capable of handling a wide range of tasks, their evaluation is becoming as challenging as their development. Most generation benchmarks currently assess LMs using abstract evaluation criteria like helpfulness and harmlessness, which often lack the flexibility and granularity of human assessment. Additionally, these benchmarks tend to focus disproportionately on specific capabilities such as instruction following, leading to coverage bias. To overcome these limitations, we introduce the BiGGen Bench, a principled generation benchmark designed to thoroughly evaluate nine distinct capabilities of LMs across 77 diverse tasks. A key feature of the BiGGen Bench is its use of instance-specific evaluation criteria, closely mirroring the nuanced discernment of human evaluation. We apply this benchmark to assess 103 frontier LMs using five evaluator LMs. Our code, data, and evaluation results are all publicly available at https://github.com/prometheus-eval/prometheus-eval/tree/main/BiGGen-Bench.
title The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
topic Computation and Language
url https://arxiv.org/abs/2406.05761