On Benchmarking Human-Like Intelligence in Machines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ying, Lance, Collins, Katherine M., Wong, Lionel, Sucholutsky, Ilia, Liu, Ryan, Weller, Adrian, Shu, Tianmin, Griffiths, Thomas L., Tenenbaum, Joshua B.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912251609350144
author Ying, Lance
Collins, Katherine M.
Wong, Lionel
Sucholutsky, Ilia
Liu, Ryan
Weller, Adrian
Shu, Tianmin
Griffiths, Thomas L.
Tenenbaum, Joshua B.
author_facet Ying, Lance
Collins, Katherine M.
Wong, Lionel
Sucholutsky, Ilia
Liu, Ryan
Weller, Adrian
Shu, Tianmin
Griffiths, Thomas L.
Tenenbaum, Joshua B.
contents Recent benchmark studies have claimed that AI has approached or even surpassed human-level performances on various cognitive tasks. However, this position paper argues that current AI evaluation paradigms are insufficient for assessing human-like cognitive capabilities. We identify a set of key shortcomings: a lack of human-validated labels, inadequate representation of human response variability and uncertainty, and reliance on simplified and ecologically-invalid tasks. We support our claims by conducting a human evaluation study on ten existing AI benchmarks, suggesting significant biases and flaws in task and label designs. To address these limitations, we propose five concrete recommendations for developing future benchmarks that will enable more rigorous and meaningful evaluations of human-like cognitive capacities in AI with various implications for such AI applications.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Benchmarking Human-Like Intelligence in Machines
Ying, Lance
Collins, Katherine M.
Wong, Lionel
Sucholutsky, Ilia
Liu, Ryan
Weller, Adrian
Shu, Tianmin
Griffiths, Thomas L.
Tenenbaum, Joshua B.
Artificial Intelligence
Recent benchmark studies have claimed that AI has approached or even surpassed human-level performances on various cognitive tasks. However, this position paper argues that current AI evaluation paradigms are insufficient for assessing human-like cognitive capabilities. We identify a set of key shortcomings: a lack of human-validated labels, inadequate representation of human response variability and uncertainty, and reliance on simplified and ecologically-invalid tasks. We support our claims by conducting a human evaluation study on ten existing AI benchmarks, suggesting significant biases and flaws in task and label designs. To address these limitations, we propose five concrete recommendations for developing future benchmarks that will enable more rigorous and meaningful evaluations of human-like cognitive capacities in AI with various implications for such AI applications.
title On Benchmarking Human-Like Intelligence in Machines
topic Artificial Intelligence
url https://arxiv.org/abs/2502.20502