On Benchmarking Human-Like Intelligence in Machines

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ying, Lance, Collins, Katherine M., Wong, Lionel, Sucholutsky, Ilia, Liu, Ryan, Weller, Adrian, Shu, Tianmin, Griffiths, Thomas L., Tenenbaum, Joshua B.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912251609350144
author Ying, Lance
Collins, Katherine M.
Wong, Lionel
Sucholutsky, Ilia
Liu, Ryan
Weller, Adrian
Shu, Tianmin
Griffiths, Thomas L.
Tenenbaum, Joshua B.
author_facet Ying, Lance
Collins, Katherine M.
Wong, Lionel
Sucholutsky, Ilia
Liu, Ryan
Weller, Adrian
Shu, Tianmin
Griffiths, Thomas L.
Tenenbaum, Joshua B.
contents Recent benchmark studies have claimed that AI has approached or even surpassed human-level performances on various cognitive tasks. However, this position paper argues that current AI evaluation paradigms are insufficient for assessing human-like cognitive capabilities. We identify a set of key shortcomings: a lack of human-validated labels, inadequate representation of human response variability and uncertainty, and reliance on simplified and ecologically-invalid tasks. We support our claims by conducting a human evaluation study on ten existing AI benchmarks, suggesting significant biases and flaws in task and label designs. To address these limitations, we propose five concrete recommendations for developing future benchmarks that will enable more rigorous and meaningful evaluations of human-like cognitive capacities in AI with various implications for such AI applications.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Benchmarking Human-Like Intelligence in Machines
Ying, Lance
Collins, Katherine M.
Wong, Lionel
Sucholutsky, Ilia
Liu, Ryan
Weller, Adrian
Shu, Tianmin
Griffiths, Thomas L.
Tenenbaum, Joshua B.
Artificial Intelligence
Recent benchmark studies have claimed that AI has approached or even surpassed human-level performances on various cognitive tasks. However, this position paper argues that current AI evaluation paradigms are insufficient for assessing human-like cognitive capabilities. We identify a set of key shortcomings: a lack of human-validated labels, inadequate representation of human response variability and uncertainty, and reliance on simplified and ecologically-invalid tasks. We support our claims by conducting a human evaluation study on ten existing AI benchmarks, suggesting significant biases and flaws in task and label designs. To address these limitations, we propose five concrete recommendations for developing future benchmarks that will enable more rigorous and meaningful evaluations of human-like cognitive capacities in AI with various implications for such AI applications.
title On Benchmarking Human-Like Intelligence in Machines
topic Artificial Intelligence
url https://arxiv.org/abs/2502.20502