Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Junlin, Tong, Shengbang, Fan, David, Ren, Yufan, Sinha, Koustuv, Torr, Philip, Kokkinos, Filippos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909816896618496
author Han, Junlin
Tong, Shengbang
Fan, David
Ren, Yufan
Sinha, Koustuv
Torr, Philip
Kokkinos, Filippos
author_facet Han, Junlin
Tong, Shengbang
Fan, David
Ren, Yufan
Sinha, Koustuv
Torr, Philip
Kokkinos, Filippos
contents Large Language Models (LLMs), despite being trained on text alone, surprisingly develop rich visual priors. These priors allow latent visual capabilities to be unlocked for vision tasks with a relatively small amount of multimodal data, and in some cases, to perform visual tasks without ever having seen an image. Through systematic analysis, we reveal that visual priors-the implicit, emergent knowledge about the visual world acquired during language pre-training-are composed of separable perception and reasoning priors with unique scaling trends and origins. We show that an LLM's latent visual reasoning ability is predominantly developed by pre-training on reasoning-centric data (e.g., code, math, academia) and scales progressively. This reasoning prior acquired from language pre-training is transferable and universally applicable to visual reasoning. In contrast, a perception prior emerges more diffusely from broad corpora, and perception ability is more sensitive to the vision encoder and visual instruction tuning data. In parallel, text describing the visual world proves crucial, though its performance impact saturates rapidly. Leveraging these insights, we propose a data-centric recipe for pre-training vision-aware LLMs and verify it in 1T token scale pre-training. Our findings are grounded in over 100 controlled experiments consuming 500,000 GPU-hours, spanning the full MLLM construction pipeline-from LLM pre-training to visual alignment and supervised multimodal fine-tuning-across five model scales, a wide range of data categories and mixtures, and multiple adaptation setups. Along with our main findings, we propose and investigate several hypotheses, and introduce the Multi-Level Existence Bench (MLE-Bench). Together, this work provides a new way of deliberately cultivating visual priors from language pre-training, paving the way for the next generation of multimodal LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
Han, Junlin
Tong, Shengbang
Fan, David
Ren, Yufan
Sinha, Koustuv
Torr, Philip
Kokkinos, Filippos
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
Large Language Models (LLMs), despite being trained on text alone, surprisingly develop rich visual priors. These priors allow latent visual capabilities to be unlocked for vision tasks with a relatively small amount of multimodal data, and in some cases, to perform visual tasks without ever having seen an image. Through systematic analysis, we reveal that visual priors-the implicit, emergent knowledge about the visual world acquired during language pre-training-are composed of separable perception and reasoning priors with unique scaling trends and origins. We show that an LLM's latent visual reasoning ability is predominantly developed by pre-training on reasoning-centric data (e.g., code, math, academia) and scales progressively. This reasoning prior acquired from language pre-training is transferable and universally applicable to visual reasoning. In contrast, a perception prior emerges more diffusely from broad corpora, and perception ability is more sensitive to the vision encoder and visual instruction tuning data. In parallel, text describing the visual world proves crucial, though its performance impact saturates rapidly. Leveraging these insights, we propose a data-centric recipe for pre-training vision-aware LLMs and verify it in 1T token scale pre-training. Our findings are grounded in over 100 controlled experiments consuming 500,000 GPU-hours, spanning the full MLLM construction pipeline-from LLM pre-training to visual alignment and supervised multimodal fine-tuning-across five model scales, a wide range of data categories and mixtures, and multiple adaptation setups. Along with our main findings, we propose and investigate several hypotheses, and introduce the Multi-Level Existence Bench (MLE-Bench). Together, this work provides a new way of deliberately cultivating visual priors from language pre-training, paving the way for the next generation of multimodal LLMs.
title Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2509.26625