BabyVision: Visual Reasoning Beyond Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Liang, Xie, Weichu, Liang, Yiyan, He, Hongfeng, Zhao, Hans, Yang, Zhibo, Huang, Zhiqi, Wu, Haoning, Lu, Haoyu, charles, Y., Bao, Yiping, Fan, Yuantao, Li, Guopeng, Shen, Haiyang, Chen, Xuanzhong, Xu, Wendong, Si, Shuzheng, Cai, Zefan, Chai, Wenhao, Huang, Ziqi, Liu, Fangfu, Liu, Tianyu, Chang, Baobao, Hu, Xiaobo, Chen, Kaiyuan, Ren, Yixin, Liu, Yang, Gong, Yuan, Li, Kuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918281295691776
author Chen, Liang
Xie, Weichu
Liang, Yiyan
He, Hongfeng
Zhao, Hans
Yang, Zhibo
Huang, Zhiqi
Wu, Haoning
Lu, Haoyu
charles, Y.
Bao, Yiping
Fan, Yuantao
Li, Guopeng
Shen, Haiyang
Chen, Xuanzhong
Xu, Wendong
Si, Shuzheng
Cai, Zefan
Chai, Wenhao
Huang, Ziqi
Liu, Fangfu
Liu, Tianyu
Chang, Baobao
Hu, Xiaobo
Chen, Kaiyuan
Ren, Yixin
Liu, Yang
Gong, Yuan
Li, Kuan
author_facet Chen, Liang
Xie, Weichu
Liang, Yiyan
He, Hongfeng
Zhao, Hans
Yang, Zhibo
Huang, Zhiqi
Wu, Haoning
Lu, Haoyu
charles, Y.
Bao, Yiping
Fan, Yuantao
Li, Guopeng
Shen, Haiyang
Chen, Xuanzhong
Xu, Wendong
Si, Shuzheng
Cai, Zefan
Chai, Wenhao
Huang, Ziqi
Liu, Fangfu
Liu, Tianyu
Chang, Baobao
Hu, Xiaobo
Chen, Kaiyuan
Ren, Yixin
Liu, Yang
Gong, Yuan
Li, Kuan
contents While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat-AI/BabyVision for reproduction.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06521
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BabyVision: Visual Reasoning Beyond Language
Chen, Liang
Xie, Weichu
Liang, Yiyan
He, Hongfeng
Zhao, Hans
Yang, Zhibo
Huang, Zhiqi
Wu, Haoning
Lu, Haoyu
charles, Y.
Bao, Yiping
Fan, Yuantao
Li, Guopeng
Shen, Haiyang
Chen, Xuanzhong
Xu, Wendong
Si, Shuzheng
Cai, Zefan
Chai, Wenhao
Huang, Ziqi
Liu, Fangfu
Liu, Tianyu
Chang, Baobao
Hu, Xiaobo
Chen, Kaiyuan
Ren, Yixin
Liu, Yang
Gong, Yuan
Li, Kuan
Computer Vision and Pattern Recognition
Computation and Language
While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat-AI/BabyVision for reproduction.
title BabyVision: Visual Reasoning Beyond Language
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2601.06521