Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Klein, Thomas, Meyen, Sascha, Brendel, Wieland, Wichmann, Felix A., Meding, Kristof
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917067696898048
author Klein, Thomas
Meyen, Sascha
Brendel, Wieland
Wichmann, Felix A.
Meding, Kristof
author_facet Klein, Thomas
Meyen, Sascha
Brendel, Wieland
Wichmann, Felix A.
Meding, Kristof
contents Benchmarking models is a key factor for the rapid progress in machine learning (ML) research. Thus, further progress depends on improving benchmarking metrics. A standard metric to measure the behavioral alignment between ML models and human observers is error consistency (EC). EC allows for more fine-grained comparisons of behavior than other metrics such as accuracy, and has been used in the influential Brain-Score benchmark to rank different DNNs by their behavioral consistency with humans. Previously, EC values have been reported without confidence intervals. However, empirically measured EC values are typically noisy -- thus, without confidence intervals, valid benchmarking conclusions are problematic. Here we improve on standard EC in two ways: First, we show how to obtain confidence intervals for EC using a bootstrapping technique, allowing us to derive significance tests for EC. Second, we propose a new computational model relating the EC between two classifiers to the implicit probability that one of them copies responses from the other. This view of EC allows us to give practical guidance to scientists regarding the number of trials required for sufficiently powerful, conclusive experiments. Finally, we use our methodology to revisit popular NeuroAI-results. We find that while the general trend of behavioral differences between humans and machines holds up to scrutiny, many reported differences between deep vision models are statistically insignificant. Our methodology enables researchers to design adequately powered experiments that can reliably detect behavioral differences between models, providing a foundation for more rigorous benchmarking of behavioral alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06645
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers
Klein, Thomas
Meyen, Sascha
Brendel, Wieland
Wichmann, Felix A.
Meding, Kristof
Neurons and Cognition
Benchmarking models is a key factor for the rapid progress in machine learning (ML) research. Thus, further progress depends on improving benchmarking metrics. A standard metric to measure the behavioral alignment between ML models and human observers is error consistency (EC). EC allows for more fine-grained comparisons of behavior than other metrics such as accuracy, and has been used in the influential Brain-Score benchmark to rank different DNNs by their behavioral consistency with humans. Previously, EC values have been reported without confidence intervals. However, empirically measured EC values are typically noisy -- thus, without confidence intervals, valid benchmarking conclusions are problematic. Here we improve on standard EC in two ways: First, we show how to obtain confidence intervals for EC using a bootstrapping technique, allowing us to derive significance tests for EC. Second, we propose a new computational model relating the EC between two classifiers to the implicit probability that one of them copies responses from the other. This view of EC allows us to give practical guidance to scientists regarding the number of trials required for sufficiently powerful, conclusive experiments. Finally, we use our methodology to revisit popular NeuroAI-results. We find that while the general trend of behavioral differences between humans and machines holds up to scrutiny, many reported differences between deep vision models are statistically insignificant. Our methodology enables researchers to design adequately powered experiments that can reliably detect behavioral differences between models, providing a foundation for more rigorous benchmarking of behavioral alignment.
title Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers
topic Neurons and Cognition
url https://arxiv.org/abs/2507.06645