Probing Multimodal Fusion in the Brain: The Dominance of Audiovisual Streams in Naturalistic Encoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abdollahi, Hamid, Majoumerd, Amir Hossein Mansouri, Baboukani, Amir Hossein Bagheri, Suratgar, Amir Abolfazl, Menhaj, Mohammad Bagher
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911077356273664
author Abdollahi, Hamid
Majoumerd, Amir Hossein Mansouri
Baboukani, Amir Hossein Bagheri
Suratgar, Amir Abolfazl
Menhaj, Mohammad Bagher
author_facet Abdollahi, Hamid
Majoumerd, Amir Hossein Mansouri
Baboukani, Amir Hossein Bagheri
Suratgar, Amir Abolfazl
Menhaj, Mohammad Bagher
contents Predicting brain activity in response to naturalistic, multimodal stimuli is a key challenge in computational neuroscience. While encoding models are becoming more powerful, their ability to generalize to truly novel contexts remains a critical, often untested, question. In this work, we developed brain encoding models using state-of-the-art visual (X-CLIP) and auditory (Whisper) feature extractors and rigorously evaluated them on both in-distribution (ID) and diverse out-of-distribution (OOD) data. Our results reveal a fundamental trade-off between model complexity and generalization: a higher-capacity attention-based model excelled on ID data, but a simpler linear model was more robust, outperforming a competitive baseline by 18\% on the OOD set. Intriguingly, we found that linguistic features did not improve predictive accuracy, suggesting that for familiar languages, neural encoding may be dominated by the continuous visual and auditory streams over redundant textual information. Spatially, our approach showed marked performance gains in the auditory cortex, underscoring the benefit of high-fidelity speech representations. Collectively, our findings demonstrate that rigorous OOD testing is essential for building robust neuro-AI models and provides nuanced insights into how model architecture, stimulus characteristics, and sensory hierarchies shape the neural encoding of our rich, multimodal world.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19052
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probing Multimodal Fusion in the Brain: The Dominance of Audiovisual Streams in Naturalistic Encoding
Abdollahi, Hamid
Majoumerd, Amir Hossein Mansouri
Baboukani, Amir Hossein Bagheri
Suratgar, Amir Abolfazl
Menhaj, Mohammad Bagher
Computer Vision and Pattern Recognition
Predicting brain activity in response to naturalistic, multimodal stimuli is a key challenge in computational neuroscience. While encoding models are becoming more powerful, their ability to generalize to truly novel contexts remains a critical, often untested, question. In this work, we developed brain encoding models using state-of-the-art visual (X-CLIP) and auditory (Whisper) feature extractors and rigorously evaluated them on both in-distribution (ID) and diverse out-of-distribution (OOD) data. Our results reveal a fundamental trade-off between model complexity and generalization: a higher-capacity attention-based model excelled on ID data, but a simpler linear model was more robust, outperforming a competitive baseline by 18\% on the OOD set. Intriguingly, we found that linguistic features did not improve predictive accuracy, suggesting that for familiar languages, neural encoding may be dominated by the continuous visual and auditory streams over redundant textual information. Spatially, our approach showed marked performance gains in the auditory cortex, underscoring the benefit of high-fidelity speech representations. Collectively, our findings demonstrate that rigorous OOD testing is essential for building robust neuro-AI models and provides nuanced insights into how model architecture, stimulus characteristics, and sensory hierarchies shape the neural encoding of our rich, multimodal world.
title Probing Multimodal Fusion in the Brain: The Dominance of Audiovisual Streams in Naturalistic Encoding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.19052