The time course of visuo-semantic representations in the human brain is captured by combining vision and language models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rong, Boyan, Gifford, Alessandro Thomas, Düzel, Emrah, Cichy, Radoslaw Martin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909658945421312
author Rong, Boyan
Gifford, Alessandro Thomas
Düzel, Emrah
Cichy, Radoslaw Martin
author_facet Rong, Boyan
Gifford, Alessandro Thomas
Düzel, Emrah
Cichy, Radoslaw Martin
contents The human visual system provides us with a rich and meaningful percept of the world, transforming retinal signals into visuo-semantic representations. For a model of these representations, here we leveraged a combination of two currently dominating approaches: vision deep neural networks (DNNs) and large language models (LLMs). Using large-scale human electroencephalography (EEG) data recorded during object image viewing, we built encoding models to predict EEG responses using representations from a vision DNN, an LLM, and their fusion. We show that the fusion encoding model outperforms encoding models based on either the vision DNN or the LLM alone, as well as previous modelling approaches, in predicting neural responses to visual stimulation. The vision DNN and the LLM complemented each other in explaining stimulus-related signal in the EEG responses. The vision DNN uniquely captured earlier and broadband EEG signals, whereas the LLM uniquely captured later and low frequency signals, as well as detailed visuo-semantic stimulus information. Together, this provides a more accurate model of the time course of visuo-semantic processing in the human brain.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19497
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The time course of visuo-semantic representations in the human brain is captured by combining vision and language models
Rong, Boyan
Gifford, Alessandro Thomas
Düzel, Emrah
Cichy, Radoslaw Martin
Neurons and Cognition
The human visual system provides us with a rich and meaningful percept of the world, transforming retinal signals into visuo-semantic representations. For a model of these representations, here we leveraged a combination of two currently dominating approaches: vision deep neural networks (DNNs) and large language models (LLMs). Using large-scale human electroencephalography (EEG) data recorded during object image viewing, we built encoding models to predict EEG responses using representations from a vision DNN, an LLM, and their fusion. We show that the fusion encoding model outperforms encoding models based on either the vision DNN or the LLM alone, as well as previous modelling approaches, in predicting neural responses to visual stimulation. The vision DNN and the LLM complemented each other in explaining stimulus-related signal in the EEG responses. The vision DNN uniquely captured earlier and broadband EEG signals, whereas the LLM uniquely captured later and low frequency signals, as well as detailed visuo-semantic stimulus information. Together, this provides a more accurate model of the time course of visuo-semantic processing in the human brain.
title The time course of visuo-semantic representations in the human brain is captured by combining vision and language models
topic Neurons and Cognition
url https://arxiv.org/abs/2506.19497