EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Dongyan, Rust, Phillip, Corrales, Angel Villar, Tan, Alvin W. M., Luthra, Mahi, Saint-James, Charles-Éric, Moritz, Rashel, Krogh-Jespersen, Sheila, Stark, Vanessa, Parimi, Surya, Shen, Jiayi, Benchekroun, Youssef, Higuchi, Yosuke, Gleize, Martin, Fizycki, Tom, Hamilakis, Nicolas, Khentout, Manel, Tsuji, Sho, Kégl, Balázs, Pino, Juan, Frank, Michael C., Dupoux, Emmanuel
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913144109006848
author Lin, Dongyan
Rust, Phillip
Corrales, Angel Villar
Tan, Alvin W. M.
Luthra, Mahi
Saint-James, Charles-Éric
Moritz, Rashel
Krogh-Jespersen, Sheila
Stark, Vanessa
Parimi, Surya
Shen, Jiayi
Benchekroun, Youssef
Higuchi, Yosuke
Gleize, Martin
Fizycki, Tom
Hamilakis, Nicolas
Khentout, Manel
Tsuji, Sho
Kégl, Balázs
Pino, Juan
Frank, Michael C.
Dupoux, Emmanuel
author_facet Lin, Dongyan
Rust, Phillip
Corrales, Angel Villar
Tan, Alvin W. M.
Luthra, Mahi
Saint-James, Charles-Éric
Moritz, Rashel
Krogh-Jespersen, Sheila
Stark, Vanessa
Parimi, Surya
Shen, Jiayi
Benchekroun, Youssef
Higuchi, Yosuke
Gleize, Martin
Fizycki, Tom
Hamilakis, Nicolas
Khentout, Manel
Tsuji, Sho
Kégl, Balázs
Pino, Juan
Frank, Michael C.
Dupoux, Emmanuel
contents Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal models. Recent research suggests current vision-language models (VLMs) trained on curated web data fail to generalize to the sparse, weakly-aligned egocentric streams produced by wearable devices, embodied agents, and infant head-cams -- and no fixed evaluation pipeline exists for measuring progress on this regime. We train VLMs on datasets with varying degrees of semantic alignment between visual and linguistic inputs, including naturalistic infant and adult egocentric videos, and evaluate them with a comprehensive suite spanning multimodal language grounding and unimodal vision and language tasks. At the core of this suite is Machine-DevBench, a corpus-grounded benchmark of lexical and grammatical competence, automatically generated from the model's training vocabulary across logarithmic frequency bins to eliminate the train/eval mismatch and low statistical power of prior developmental benchmarks. Our results show that current VLM paradigms hinge on the tight semantic alignment of curated data and fail to exploit the weakly-aligned signal that dominates naturalistic egocentric input -- the very regime in which humans thrive. To motivate progress, we introduce the EgoBabyVLM Challenge to drive the development of models capable of grounded language learning from the kind of naturalistic data that human infants experience.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19130
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
Lin, Dongyan
Rust, Phillip
Corrales, Angel Villar
Tan, Alvin W. M.
Luthra, Mahi
Saint-James, Charles-Éric
Moritz, Rashel
Krogh-Jespersen, Sheila
Stark, Vanessa
Parimi, Surya
Shen, Jiayi
Benchekroun, Youssef
Higuchi, Yosuke
Gleize, Martin
Fizycki, Tom
Hamilakis, Nicolas
Khentout, Manel
Tsuji, Sho
Kégl, Balázs
Pino, Juan
Frank, Michael C.
Dupoux, Emmanuel
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal models. Recent research suggests current vision-language models (VLMs) trained on curated web data fail to generalize to the sparse, weakly-aligned egocentric streams produced by wearable devices, embodied agents, and infant head-cams -- and no fixed evaluation pipeline exists for measuring progress on this regime. We train VLMs on datasets with varying degrees of semantic alignment between visual and linguistic inputs, including naturalistic infant and adult egocentric videos, and evaluate them with a comprehensive suite spanning multimodal language grounding and unimodal vision and language tasks. At the core of this suite is Machine-DevBench, a corpus-grounded benchmark of lexical and grammatical competence, automatically generated from the model's training vocabulary across logarithmic frequency bins to eliminate the train/eval mismatch and low statistical power of prior developmental benchmarks. Our results show that current VLM paradigms hinge on the tight semantic alignment of curated data and fail to exploit the weakly-aligned signal that dominates naturalistic egocentric input -- the very regime in which humans thrive. To motivate progress, we introduce the EgoBabyVLM Challenge to drive the development of models capable of grounded language learning from the kind of naturalistic data that human infants experience.
title EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.19130