Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Jianing, Li, Runan, Pang, Honglin, Xia, Ding, Zhu, Zhou, Zhang, Qian, Li, Chuntao, Yang, Xi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908947356581888
author Zhang, Jianing
Li, Runan
Pang, Honglin
Xia, Ding
Zhu, Zhou
Zhang, Qian
Li, Chuntao
Yang, Xi
author_facet Zhang, Jianing
Li, Runan
Pang, Honglin
Xia, Ding
Zhu, Zhou
Zhang, Qian
Li, Chuntao
Yang, Xi
contents Deciphering ancient Chinese Oracle Bone Script (OBS) is a challenging task that offers insights into the beliefs, systems, and culture of the ancient era. Existing approaches treat decipherment as a closed-set image recognition problem, which fails to bridge the ``interpretation gap'': while individual characters are often unique and rare, they are composed of a limited set of recurring, pictographic components that carry transferable semantic meanings. To leverage this structural logic, we propose an agent-driven Vision-Language Model (VLM) framework that integrates a VLM for precise visual grounding with an LLM-based agent to automate a reasoning chain of component identification, graph-based knowledge retrieval, and relationship inference for linguistically accurate interpretation. To support this, we also introduce OB-Radix, an expert-annotated dataset providing structural and semantic data absent from prior corpora, comprising 1,022 character images (934 unique characters) and 1,853 fine-grained component images across 478 distinct components with verified explanations. By evaluating our system across three benchmarks of different tasks, we demonstrate that our framework yields more detailed and precise decipherments compared to baseline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06711
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation
Zhang, Jianing
Li, Runan
Pang, Honglin
Xia, Ding
Zhu, Zhou
Zhang, Qian
Li, Chuntao
Yang, Xi
Computer Vision and Pattern Recognition
Computation and Language
Deciphering ancient Chinese Oracle Bone Script (OBS) is a challenging task that offers insights into the beliefs, systems, and culture of the ancient era. Existing approaches treat decipherment as a closed-set image recognition problem, which fails to bridge the ``interpretation gap'': while individual characters are often unique and rare, they are composed of a limited set of recurring, pictographic components that carry transferable semantic meanings. To leverage this structural logic, we propose an agent-driven Vision-Language Model (VLM) framework that integrates a VLM for precise visual grounding with an LLM-based agent to automate a reasoning chain of component identification, graph-based knowledge retrieval, and relationship inference for linguistically accurate interpretation. To support this, we also introduce OB-Radix, an expert-annotated dataset providing structural and semantic data absent from prior corpora, comprising 1,022 character images (934 unique characters) and 1,853 fine-grained component images across 478 distinct components with verified explanations. By evaluating our system across three benchmarks of different tasks, we demonstrate that our framework yields more detailed and precise decipherments compared to baseline methods.
title Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2604.06711