VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jinzhou, Ma, Zhengwu, Li, Jixing, Tang, Baoping, Lu, Zitong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914609119625216
author Wu, Jinzhou
Ma, Zhengwu
Li, Jixing
Tang, Baoping
Lu, Zitong
author_facet Wu, Jinzhou
Ma, Zhengwu
Li, Jixing
Tang, Baoping
Lu, Zitong
contents Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-language learning makes text representations more human-like during natural reading. Here, we address this question by comparing tightly matched LLM and vision-language model (VLM) pairs under a strictly text-only setting, allowing us to isolate the effect of multimodal training history from online visual input or cross-modal fusion. We evaluate model alignment with a human natural-reading dataset that includes whole-cortex fMRI responses and synchronized eye-tracking saccades. Our findings demonstrate that multimodal pretraining may not confer a uniform, global advantage in human alignment during natural reading, indicating that language-internal representations remain the key factor for modeling human text processing. However, the VLM advantage could emerge more selectively when sentences contain stronger visual semantic content, with converging evidence from both fMRI and eye-movement alignments. Together, our findings provide a controlled in silico framework for testing how visual learning history shapes model-human alignment of language processing, suggesting that multimodal pretraining contributes selectively rather than globally to human-like language representations during natural reading.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28818
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading
Wu, Jinzhou
Ma, Zhengwu
Li, Jixing
Tang, Baoping
Lu, Zitong
Computation and Language
Neurons and Cognition
Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-language learning makes text representations more human-like during natural reading. Here, we address this question by comparing tightly matched LLM and vision-language model (VLM) pairs under a strictly text-only setting, allowing us to isolate the effect of multimodal training history from online visual input or cross-modal fusion. We evaluate model alignment with a human natural-reading dataset that includes whole-cortex fMRI responses and synchronized eye-tracking saccades. Our findings demonstrate that multimodal pretraining may not confer a uniform, global advantage in human alignment during natural reading, indicating that language-internal representations remain the key factor for modeling human text processing. However, the VLM advantage could emerge more selectively when sentences contain stronger visual semantic content, with converging evidence from both fMRI and eye-movement alignments. Together, our findings provide a controlled in silico framework for testing how visual learning history shapes model-human alignment of language processing, suggesting that multimodal pretraining contributes selectively rather than globally to human-like language representations during natural reading.
title VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading
topic Computation and Language
Neurons and Cognition
url https://arxiv.org/abs/2605.28818