TextClass Benchmark: A Continuous Elo Rating of LLMs in Social Sciences

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: González-Bustamante, Bastián
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909419032281088
author González-Bustamante, Bastián
author_facet González-Bustamante, Bastián
contents The TextClass Benchmark project is an ongoing, continuous benchmarking process that aims to provide a comprehensive, fair, and dynamic evaluation of LLMs and transformers for text classification tasks. This evaluation spans various domains and languages in social sciences disciplines engaged in NLP and text-as-data approach. The leaderboards present performance metrics and relative ranking using a tailored Elo rating system. With each leaderboard cycle, novel models are added, fixed test sets can be replaced for unseen, equivalent data to test generalisation power, ratings are updated, and a Meta-Elo leaderboard combines and weights domain-specific leaderboards. This article presents the rationale and motivation behind the project, explains the Elo rating system in detail, and estimates Meta-Elo across different classification tasks in social science disciplines. We also present a snapshot of the first cycle of classification tasks on incivility data in Chinese, English, German and Russian. This ongoing benchmarking process includes not only additional languages such as Arabic, Hindi, and Spanish but also a classification of policy agenda topics, misinformation, among others.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00539
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TextClass Benchmark: A Continuous Elo Rating of LLMs in Social Sciences
González-Bustamante, Bastián
Computation and Language
Artificial Intelligence
68T50 (Primary) 91F10, 91F20 (Secondary)
The TextClass Benchmark project is an ongoing, continuous benchmarking process that aims to provide a comprehensive, fair, and dynamic evaluation of LLMs and transformers for text classification tasks. This evaluation spans various domains and languages in social sciences disciplines engaged in NLP and text-as-data approach. The leaderboards present performance metrics and relative ranking using a tailored Elo rating system. With each leaderboard cycle, novel models are added, fixed test sets can be replaced for unseen, equivalent data to test generalisation power, ratings are updated, and a Meta-Elo leaderboard combines and weights domain-specific leaderboards. This article presents the rationale and motivation behind the project, explains the Elo rating system in detail, and estimates Meta-Elo across different classification tasks in social science disciplines. We also present a snapshot of the first cycle of classification tasks on incivility data in Chinese, English, German and Russian. This ongoing benchmarking process includes not only additional languages such as Arabic, Hindi, and Spanish but also a classification of policy agenda topics, misinformation, among others.
title TextClass Benchmark: A Continuous Elo Rating of LLMs in Social Sciences
topic Computation and Language
Artificial Intelligence
68T50 (Primary) 91F10, 91F20 (Secondary)
url https://arxiv.org/abs/2412.00539