Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Zekai, Cai, Zhixi, Stefanov, Kalin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917080814583808
author Shi, Zekai
Cai, Zhixi
Stefanov, Kalin
author_facet Shi, Zekai
Cai, Zhixi
Stefanov, Kalin
contents Typically, children start to learn their first words between 6 and 9 months, linking spoken utterances to their visual referents. Without prior knowledge, a word encountered for the first time can be interpreted in countless ways; it might refer to any of the objects in the environment, their components, or attributes. Using longitudinal, egocentric, and ecologically valid data from the experience of one child, in this work, we propose a self-supervised and biologically plausible strategy to learn strong visual representations. Our masked autoencoder-based visual backbone incorporates knowledge about the blind spot in human eyes to define a novel masking strategy. This mask and reconstruct approach attempts to mimic the way the human brain fills the gaps in the eyes' field of view. This represents a significant shift from standard random masking strategies, which are difficult to justify from a biological perspective. The pretrained encoder is utilized in a contrastive learning-based video-text model capable of acquiring word-referent mappings. Extensive evaluation suggests that the proposed biologically plausible masking strategy is at least as effective as random masking for learning word-referent mappings from cross-situational and temporally extended episodes.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11725
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video
Shi, Zekai
Cai, Zhixi
Stefanov, Kalin
Computer Vision and Pattern Recognition
Artificial Intelligence
Typically, children start to learn their first words between 6 and 9 months, linking spoken utterances to their visual referents. Without prior knowledge, a word encountered for the first time can be interpreted in countless ways; it might refer to any of the objects in the environment, their components, or attributes. Using longitudinal, egocentric, and ecologically valid data from the experience of one child, in this work, we propose a self-supervised and biologically plausible strategy to learn strong visual representations. Our masked autoencoder-based visual backbone incorporates knowledge about the blind spot in human eyes to define a novel masking strategy. This mask and reconstruct approach attempts to mimic the way the human brain fills the gaps in the eyes' field of view. This represents a significant shift from standard random masking strategies, which are difficult to justify from a biological perspective. The pretrained encoder is utilized in a contrastive learning-based video-text model capable of acquiring word-referent mappings. Extensive evaluation suggests that the proposed biologically plausible masking strategy is at least as effective as random masking for learning word-referent mappings from cross-situational and temporally extended episodes.
title Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.11725