Stacked Regression using Off-the-shelf, Stimulus-tuned and Fine-tuned Neural Networks for Predicting fMRI Brain Responses to Movies (Algonauts 2025 Report)

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Scholz, Robert, Bagga, Kunal, Ahrends, Christine, Barbano, Carlo Alberto
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918156232032256
author Scholz, Robert
Bagga, Kunal
Ahrends, Christine
Barbano, Carlo Alberto
author_facet Scholz, Robert
Bagga, Kunal
Ahrends, Christine
Barbano, Carlo Alberto
contents We present our submission to the Algonauts 2025 Challenge, where the goal is to predict fMRI brain responses to movie stimuli. Our approach integrates multimodal representations from large language models, video encoders, audio models, and vision-language models, combining both off-the-shelf and fine-tuned variants. To improve performance, we enhanced textual inputs with detailed transcripts and summaries, and we explored stimulus-tuning and fine-tuning strategies for language and vision models. Predictions from individual models were combined using stacked regression, yielding solid results. Our submission, under the team name Seinfeld, ranked 10th. We make all code and resources publicly available, contributing to ongoing efforts in developing multimodal encoding models for brain activity.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06235
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stacked Regression using Off-the-shelf, Stimulus-tuned and Fine-tuned Neural Networks for Predicting fMRI Brain Responses to Movies (Algonauts 2025 Report)
Scholz, Robert
Bagga, Kunal
Ahrends, Christine
Barbano, Carlo Alberto
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Neurons and Cognition
We present our submission to the Algonauts 2025 Challenge, where the goal is to predict fMRI brain responses to movie stimuli. Our approach integrates multimodal representations from large language models, video encoders, audio models, and vision-language models, combining both off-the-shelf and fine-tuned variants. To improve performance, we enhanced textual inputs with detailed transcripts and summaries, and we explored stimulus-tuning and fine-tuning strategies for language and vision models. Predictions from individual models were combined using stacked regression, yielding solid results. Our submission, under the team name Seinfeld, ranked 10th. We make all code and resources publicly available, contributing to ongoing efforts in developing multimodal encoding models for brain activity.
title Stacked Regression using Off-the-shelf, Stimulus-tuned and Fine-tuned Neural Networks for Predicting fMRI Brain Responses to Movies (Algonauts 2025 Report)
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Neurons and Cognition
url https://arxiv.org/abs/2510.06235