Towards Reliable Testing for Multiple Information Retrieval System Comparisons

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Otero, David, Parapar, Javier, Barreiro, Álvaro
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908459780276224
author Otero, David
Parapar, Javier
Barreiro, Álvaro
author_facet Otero, David
Parapar, Javier
Barreiro, Álvaro
contents Null Hypothesis Significance Testing is the \textit{de facto} tool for assessing effectiveness differences between Information Retrieval systems. Researchers use statistical tests to check whether those differences will generalise to online settings or are just due to the samples observed in the laboratory. Much work has been devoted to studying which test is the most reliable when comparing a pair of systems, but most of the IR real-world experiments involve more than two. In the multiple comparisons scenario, testing several systems simultaneously may inflate the errors committed by the tests. In this paper, we use a new approach to assess the reliability of multiple comparison procedures using simulated and real TREC data. Experiments show that Wilcoxon plus the Benjamini-Hochberg correction yields Type I error rates according to the significance level for typical sample sizes while being the best test in terms of statistical power.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03930
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Reliable Testing for Multiple Information Retrieval System Comparisons
Otero, David
Parapar, Javier
Barreiro, Álvaro
Information Retrieval
Null Hypothesis Significance Testing is the \textit{de facto} tool for assessing effectiveness differences between Information Retrieval systems. Researchers use statistical tests to check whether those differences will generalise to online settings or are just due to the samples observed in the laboratory. Much work has been devoted to studying which test is the most reliable when comparing a pair of systems, but most of the IR real-world experiments involve more than two. In the multiple comparisons scenario, testing several systems simultaneously may inflate the errors committed by the tests. In this paper, we use a new approach to assess the reliability of multiple comparison procedures using simulated and real TREC data. Experiments show that Wilcoxon plus the Benjamini-Hochberg correction yields Type I error rates according to the significance level for typical sample sizes while being the best test in terms of statistical power.
title Towards Reliable Testing for Multiple Information Retrieval System Comparisons
topic Information Retrieval
url https://arxiv.org/abs/2501.03930