Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Thorat, Sushrut, Doerig, Adrien, Kroner, Alexander, Amme, Carmen, Kietzmann, Tim C.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918203825848320
author Thorat, Sushrut
Doerig, Adrien
Kroner, Alexander
Amme, Carmen
Kietzmann, Tim C.
author_facet Thorat, Sushrut
Doerig, Adrien
Kroner, Alexander
Amme, Carmen
Kietzmann, Tim C.
contents Scenes are complex, yet structured collections of parts, including objects and surfaces, that exhibit spatial and semantic relations to one another. An effective visual system therefore needs unified scene representations that relate scene parts to their location and their co-occurrence. We hypothesize that this structure can be learned self-supervised from natural experience by exploiting the temporal regularities of active vision: each fixation reveals a locally-detailed glimpse that is statistically related to the previous one via co-occurrence and saccade-conditioned spatial regularities. We instantiate this idea with Glimpse Prediction Networks (GPNs) -- recurrent models trained to predict the feature embedding of the next glimpse along human-like scanpaths over natural scenes. GPNs successfully learn co-occurrence structure and, when given relative saccade location vectors, show sensitivity to spatial arrangement. Furthermore, recurrent variants of GPNs were able to integrate information across glimpses into a unified scene representation. Notably, these scene representations align strongly with human fMRI responses during natural-scene viewing across mid/high-level visual cortex. Critically, GPNs outperform architecture- and dataset-matched controls trained with explicit semantic objectives, and match or exceed strong modern vision baselines, leaving little unique variance for those alternatives. These results establish next-glimpse prediction during active vision as a biologically plausible, self-supervised route to brain-aligned scene representations learned from natural visual experience.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12715
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
Thorat, Sushrut
Doerig, Adrien
Kroner, Alexander
Amme, Carmen
Kietzmann, Tim C.
Neurons and Cognition
Computer Vision and Pattern Recognition
Scenes are complex, yet structured collections of parts, including objects and surfaces, that exhibit spatial and semantic relations to one another. An effective visual system therefore needs unified scene representations that relate scene parts to their location and their co-occurrence. We hypothesize that this structure can be learned self-supervised from natural experience by exploiting the temporal regularities of active vision: each fixation reveals a locally-detailed glimpse that is statistically related to the previous one via co-occurrence and saccade-conditioned spatial regularities. We instantiate this idea with Glimpse Prediction Networks (GPNs) -- recurrent models trained to predict the feature embedding of the next glimpse along human-like scanpaths over natural scenes. GPNs successfully learn co-occurrence structure and, when given relative saccade location vectors, show sensitivity to spatial arrangement. Furthermore, recurrent variants of GPNs were able to integrate information across glimpses into a unified scene representation. Notably, these scene representations align strongly with human fMRI responses during natural-scene viewing across mid/high-level visual cortex. Critically, GPNs outperform architecture- and dataset-matched controls trained with explicit semantic objectives, and match or exceed strong modern vision baselines, leaving little unique variance for those alternatives. These results establish next-glimpse prediction during active vision as a biologically plausible, self-supervised route to brain-aligned scene representations learned from natural visual experience.
title Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
topic Neurons and Cognition
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.12715