Modality-Agnostic fMRI Decoding of Vision and Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nikolaus, Mitja, Mozafari, Milad, Asher, Nicholas, Reddy, Leila, VanRullen, Rufin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929281337131008
author Nikolaus, Mitja
Mozafari, Milad
Asher, Nicholas
Reddy, Leila
VanRullen, Rufin
author_facet Nikolaus, Mitja
Mozafari, Milad
Asher, Nicholas
Reddy, Leila
VanRullen, Rufin
contents Previous studies have shown that it is possible to map brain activation data of subjects viewing images onto the feature representation space of not only vision models (modality-specific decoding) but also language models (cross-modal decoding). In this work, we introduce and use a new large-scale fMRI dataset (~8,500 trials per subject) of people watching both images and text descriptions of such images. This novel dataset enables the development of modality-agnostic decoders: a single decoder that can predict which stimulus a subject is seeing, irrespective of the modality (image or text) in which the stimulus is presented. We train and evaluate such decoders to map brain signals onto stimulus representations from a large range of publicly available vision, language and multimodal (vision+language) models. Our findings reveal that (1) modality-agnostic decoders perform as well as (and sometimes even better than) modality-specific decoders (2) modality-agnostic decoders mapping brain data onto representations from unimodal models perform as well as decoders relying on multimodal representations (3) while language and low-level visual (occipital) brain regions are best at decoding text and image stimuli, respectively, high-level visual (temporal) regions perform well on both stimulus types.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11771
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Modality-Agnostic fMRI Decoding of Vision and Language
Nikolaus, Mitja
Mozafari, Milad
Asher, Nicholas
Reddy, Leila
VanRullen, Rufin
Computer Vision and Pattern Recognition
Computation and Language
Previous studies have shown that it is possible to map brain activation data of subjects viewing images onto the feature representation space of not only vision models (modality-specific decoding) but also language models (cross-modal decoding). In this work, we introduce and use a new large-scale fMRI dataset (~8,500 trials per subject) of people watching both images and text descriptions of such images. This novel dataset enables the development of modality-agnostic decoders: a single decoder that can predict which stimulus a subject is seeing, irrespective of the modality (image or text) in which the stimulus is presented. We train and evaluate such decoders to map brain signals onto stimulus representations from a large range of publicly available vision, language and multimodal (vision+language) models. Our findings reveal that (1) modality-agnostic decoders perform as well as (and sometimes even better than) modality-specific decoders (2) modality-agnostic decoders mapping brain data onto representations from unimodal models perform as well as decoders relying on multimodal representations (3) while language and low-level visual (occipital) brain regions are best at decoding text and image stimuli, respectively, high-level visual (temporal) regions perform well on both stimulus types.
title Modality-Agnostic fMRI Decoding of Vision and Language
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2403.11771