See the Text: From Tokenization to Visual Reading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xing, Ling, Yan, Rui, Wang, Alex Jinpeng, Li, Zechao, Tang, Jinhui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918410132127744
author Xing, Ling
Yan, Rui
Wang, Alex Jinpeng
Li, Zechao
Tang, Jinhui
author_facet Xing, Ling
Yan, Rui
Wang, Alex Jinpeng
Li, Zechao
Tang, Jinhui
contents People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle typos, distorted fonts, and various scripts effectively. Modern large language models (LLMs), however, rely on subword tokenization, fragmenting text into pieces from a fixed vocabulary. While effective for high-resource languages, this approach over-segments low-resource languages, yielding long, linguistically meaningless sequences and inflating computation. In this work, we challenge this entrenched paradigm and move toward a vision-centric alternative. Our method, SeeTok, renders text as images (visual-text) and leverages pretrained multimodal LLMs to interpret them, reusing strong OCR and text-vision alignment abilities learned from large-scale multimodal training. Across three different language tasks, SeeTok matches or surpasses subword tokenizers while requiring 4.43 times fewer tokens and reducing FLOPs by 70.5%, with additional gains in cross-lingual generalization, robustness to typographic noise, and linguistic hierarchy. SeeTok signals a shift from symbolic tokenization to human-like visual reading, and takes a step toward more natural and cognitively inspired language models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18840
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle See the Text: From Tokenization to Visual Reading
Xing, Ling
Yan, Rui
Wang, Alex Jinpeng
Li, Zechao
Tang, Jinhui
Computer Vision and Pattern Recognition
Computation and Language
People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle typos, distorted fonts, and various scripts effectively. Modern large language models (LLMs), however, rely on subword tokenization, fragmenting text into pieces from a fixed vocabulary. While effective for high-resource languages, this approach over-segments low-resource languages, yielding long, linguistically meaningless sequences and inflating computation. In this work, we challenge this entrenched paradigm and move toward a vision-centric alternative. Our method, SeeTok, renders text as images (visual-text) and leverages pretrained multimodal LLMs to interpret them, reusing strong OCR and text-vision alignment abilities learned from large-scale multimodal training. Across three different language tasks, SeeTok matches or surpasses subword tokenizers while requiring 4.43 times fewer tokens and reducing FLOPs by 70.5%, with additional gains in cross-lingual generalization, robustness to typographic noise, and linguistic hierarchy. SeeTok signals a shift from symbolic tokenization to human-like visual reading, and takes a step toward more natural and cognitively inspired language models.
title See the Text: From Tokenization to Visual Reading
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.18840