Does Spatial Cognition Emerge in Frontier Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramakrishnan, Santhosh Kumar, Wijmans, Erik, Kraehenbuehl, Philipp, Koltun, Vladlen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913798919553024
author Ramakrishnan, Santhosh Kumar
Wijmans, Erik
Kraehenbuehl, Philipp
Koltun, Vladlen
author_facet Ramakrishnan, Santhosh Kumar
Wijmans, Erik
Kraehenbuehl, Philipp
Koltun, Vladlen
contents Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that are brought to bear when an organism traverses physical environments, smaller-scale reasoning about object shapes and layouts, and cognitive infrastructure such as spatial attention and memory. For many tasks, we instantiate parallel presentations via text and images, allowing us to benchmark both large language models and large multimodal models. Results suggest that contemporary frontier models fall short of the spatial intelligence of animals, performing near chance level on a number of classic tests of animal cognition. Code and data are available: https://github.com/apple/ml-space-benchmark
format Preprint
id arxiv_https___arxiv_org_abs_2410_06468
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Does Spatial Cognition Emerge in Frontier Models?
Ramakrishnan, Santhosh Kumar
Wijmans, Erik
Kraehenbuehl, Philipp
Koltun, Vladlen
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that are brought to bear when an organism traverses physical environments, smaller-scale reasoning about object shapes and layouts, and cognitive infrastructure such as spatial attention and memory. For many tasks, we instantiate parallel presentations via text and images, allowing us to benchmark both large language models and large multimodal models. Results suggest that contemporary frontier models fall short of the spatial intelligence of animals, performing near chance level on a number of classic tests of animal cognition. Code and data are available: https://github.com/apple/ml-space-benchmark
title Does Spatial Cognition Emerge in Frontier Models?
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.06468