Exposing Assumptions in AI Benchmarks through Cognitive Modelling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rystrøm, Jonathan H., Enevoldsen, Kenneth C.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917786727481344
author Rystrøm, Jonathan H.
Enevoldsen, Kenneth C.
author_facet Rystrøm, Jonathan H.
Enevoldsen, Kenneth C.
contents Cultural AI benchmarks often rely on implicit assumptions about measured constructs, leading to vague formulations with poor validity and unclear interrelations. We propose exposing these assumptions using explicit cognitive models formulated as Structural Equation Models. Using cross-lingual alignment transfer as an example, we show how this approach can answer key research questions and identify missing datasets. This framework grounds benchmark construction theoretically and guides dataset development to improve construct measurement. By embracing transparency, we move towards more rigorous, cumulative AI evaluation science, challenging researchers to critically examine their assessment foundations.
format Preprint
id arxiv_https___arxiv_org_abs_2409_16849
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exposing Assumptions in AI Benchmarks through Cognitive Modelling
Rystrøm, Jonathan H.
Enevoldsen, Kenneth C.
Artificial Intelligence
Computation and Language
Cultural AI benchmarks often rely on implicit assumptions about measured constructs, leading to vague formulations with poor validity and unclear interrelations. We propose exposing these assumptions using explicit cognitive models formulated as Structural Equation Models. Using cross-lingual alignment transfer as an example, we show how this approach can answer key research questions and identify missing datasets. This framework grounds benchmark construction theoretically and guides dataset development to improve construct measurement. By embracing transparency, we move towards more rigorous, cumulative AI evaluation science, challenging researchers to critically examine their assessment foundations.
title Exposing Assumptions in AI Benchmarks through Cognitive Modelling
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2409.16849