CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Luo, Yuxuan, Tang, Jiaqi, Huang, Chenyi, Hao, Feiyang, Lian, Zhouhui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911356027928576
author Luo, Yuxuan
Tang, Jiaqi
Huang, Chenyi
Hao, Feiyang
Lian, Zhouhui
author_facet Luo, Yuxuan
Tang, Jiaqi
Huang, Chenyi
Hao, Feiyang
Lian, Zhouhui
contents Chinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fail to contextualize their intricate scripts, because of limited annotated data and poor visual-semantic alignment. We propose CalliReader, a vision-language model (VLM) that solves the Chinese Calligraphy Contextualization (CC$^2$) problem through three innovations: (1) character-wise slicing for precise character extraction and sorting, (2) CalliAlign for visual-text token compression and alignment, (3) embedding instruction tuning (e-IT) for improving alignment and addressing data scarcity. We also build CalliBench, the first benchmark for full-page calligraphic contextualization, addressing three critical issues in previous OCR and VQA approaches: fragmented context, shallow reasoning, and hallucination. Extensive experiments including user studies have been conducted to verify our CalliReader's \textbf{superiority to other state-of-the-art methods and even human professionals in page-level calligraphy recognition and interpretation}, achieving higher accuracy while reducing hallucination. Comparisons with reasoning models highlight the importance of accurate recognition as a prerequisite for reliable comprehension. Quantitative analyses validate CalliReader's efficiency; evaluations on document and real-world benchmarks confirm its robust generalization ability.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06472
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
Luo, Yuxuan
Tang, Jiaqi
Huang, Chenyi
Hao, Feiyang
Lian, Zhouhui
Computer Vision and Pattern Recognition
Multimedia
Chinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fail to contextualize their intricate scripts, because of limited annotated data and poor visual-semantic alignment. We propose CalliReader, a vision-language model (VLM) that solves the Chinese Calligraphy Contextualization (CC$^2$) problem through three innovations: (1) character-wise slicing for precise character extraction and sorting, (2) CalliAlign for visual-text token compression and alignment, (3) embedding instruction tuning (e-IT) for improving alignment and addressing data scarcity. We also build CalliBench, the first benchmark for full-page calligraphic contextualization, addressing three critical issues in previous OCR and VQA approaches: fragmented context, shallow reasoning, and hallucination. Extensive experiments including user studies have been conducted to verify our CalliReader's \textbf{superiority to other state-of-the-art methods and even human professionals in page-level calligraphy recognition and interpretation}, achieving higher accuracy while reducing hallucination. Comparisons with reasoning models highlight the importance of accurate recognition as a prerequisite for reliable comprehension. Quantitative analyses validate CalliReader's efficiency; evaluations on document and real-world benchmarks confirm its robust generalization ability.
title CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2503.06472