BabyVision: Visual Reasoning Beyond Language
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918281295691776 |
|---|---|
| author | Chen, Liang Xie, Weichu Liang, Yiyan He, Hongfeng Zhao, Hans Yang, Zhibo Huang, Zhiqi Wu, Haoning Lu, Haoyu charles, Y. Bao, Yiping Fan, Yuantao Li, Guopeng Shen, Haiyang Chen, Xuanzhong Xu, Wendong Si, Shuzheng Cai, Zefan Chai, Wenhao Huang, Ziqi Liu, Fangfu Liu, Tianyu Chang, Baobao Hu, Xiaobo Chen, Kaiyuan Ren, Yixin Liu, Yang Gong, Yuan Li, Kuan |
| author_facet | Chen, Liang Xie, Weichu Liang, Yiyan He, Hongfeng Zhao, Hans Yang, Zhibo Huang, Zhiqi Wu, Haoning Lu, Haoyu charles, Y. Bao, Yiping Fan, Yuantao Li, Guopeng Shen, Haiyang Chen, Xuanzhong Xu, Wendong Si, Shuzheng Cai, Zefan Chai, Wenhao Huang, Ziqi Liu, Fangfu Liu, Tianyu Chang, Baobao Hu, Xiaobo Chen, Kaiyuan Ren, Yixin Liu, Yang Gong, Yuan Li, Kuan |
| contents | While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat-AI/BabyVision for reproduction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_06521 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | BabyVision: Visual Reasoning Beyond Language Chen, Liang Xie, Weichu Liang, Yiyan He, Hongfeng Zhao, Hans Yang, Zhibo Huang, Zhiqi Wu, Haoning Lu, Haoyu charles, Y. Bao, Yiping Fan, Yuantao Li, Guopeng Shen, Haiyang Chen, Xuanzhong Xu, Wendong Si, Shuzheng Cai, Zefan Chai, Wenhao Huang, Ziqi Liu, Fangfu Liu, Tianyu Chang, Baobao Hu, Xiaobo Chen, Kaiyuan Ren, Yixin Liu, Yang Gong, Yuan Li, Kuan Computer Vision and Pattern Recognition Computation and Language While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat-AI/BabyVision for reproduction. |
| title | BabyVision: Visual Reasoning Beyond Language |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2601.06521 |