Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sälevä, Jonne, Ataman, Duygu, Lignos, Constantine
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914206146625536
author Sälevä, Jonne
Ataman, Duygu
Lignos, Constantine
author_facet Sälevä, Jonne
Ataman, Duygu
Lignos, Constantine
contents We introduce a set of resampling-based methods for quantifying uncertainty and statistical precision of evaluation metrics in multilingual and/or multitask NLP benchmarks. We show how experimental variation in performance scores arises from both model and data-related sources, and that accounting for both of them is necessary to avoid substantially underestimating the overall variability over hypothetical replications. Using multilingual question answering, machine translation, and named entity recognition as example tasks, we also demonstrate how resampling methods are useful for quantifying the replication uncertainty of various quantities used in leaderboards such as model rankings and pairwise differences between models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22612
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation
Sälevä, Jonne
Ataman, Duygu
Lignos, Constantine
Computation and Language
We introduce a set of resampling-based methods for quantifying uncertainty and statistical precision of evaluation metrics in multilingual and/or multitask NLP benchmarks. We show how experimental variation in performance scores arises from both model and data-related sources, and that accounting for both of them is necessary to avoid substantially underestimating the overall variability over hypothetical replications. Using multilingual question answering, machine translation, and named entity recognition as example tasks, we also demonstrate how resampling methods are useful for quantifying the replication uncertainty of various quantities used in leaderboards such as model rankings and pairwise differences between models.
title Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation
topic Computation and Language
url https://arxiv.org/abs/2509.22612