Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: López-Cardona, Ángela, Idesis, Sebastián, Masias-Bruns, Mireia, Abadal, Sergi, Arapakis, Ioannis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911222528475136
author López-Cardona, Ángela
Idesis, Sebastián
Masias-Bruns, Mireia
Abadal, Sergi
Arapakis, Ioannis
author_facet López-Cardona, Ángela
Idesis, Sebastián
Masias-Bruns, Mireia
Abadal, Sergi
Arapakis, Ioannis
contents Do brains and language models converge toward the same internal representations of the world? Recent years have seen a rise in studies of neural activations and model alignment. In this work, we review 25 fMRI-based studies published between 2023 and 2025 and explicitly confront their findings with two key hypotheses: (i) the Platonic Representation Hypothesis -- that as models scale and improve, they converge to a representation of the real world, and (ii) the Intermediate-Layer Advantage -- that intermediate (mid-depth) layers often encode richer, more generalizable features. Our findings provide converging evidence that models and brains may share abstract representational structures, supporting both hypotheses and motivating further research on brain-model alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17833
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
López-Cardona, Ángela
Idesis, Sebastián
Masias-Bruns, Mireia
Abadal, Sergi
Arapakis, Ioannis
Neurons and Cognition
Artificial Intelligence
Do brains and language models converge toward the same internal representations of the world? Recent years have seen a rise in studies of neural activations and model alignment. In this work, we review 25 fMRI-based studies published between 2023 and 2025 and explicitly confront their findings with two key hypotheses: (i) the Platonic Representation Hypothesis -- that as models scale and improve, they converge to a representation of the real world, and (ii) the Intermediate-Layer Advantage -- that intermediate (mid-depth) layers often encode richer, more generalizable features. Our findings provide converging evidence that models and brains may share abstract representational structures, supporting both hypotheses and motivating further research on brain-model alignment.
title Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
topic Neurons and Cognition
Artificial Intelligence
url https://arxiv.org/abs/2510.17833