MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Zhang, Liu, Yuliang, Liu, Qiang, Ma, Zhiyin, Zhang, Ziyang, Zhang, Shuo, Yang, Biao, Guo, Zidun, Zhang, Jiarui, Wang, Xinyu, Bai, Xiang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908818734055424
author Li, Zhang
Liu, Yuliang
Liu, Qiang
Ma, Zhiyin
Zhang, Ziyang
Zhang, Shuo
Yang, Biao
Guo, Zidun
Zhang, Jiarui
Wang, Xinyu
Bai, Xiang
author_facet Li, Zhang
Liu, Yuliang
Liu, Qiang
Ma, Zhiyin
Zhang, Ziyang
Zhang, Shuo
Yang, Biao
Guo, Zidun
Zhang, Jiarui
Wang, Xinyu
Bai, Xiang
contents We introduce MonkeyOCR, a document parsing model that advances the state of the art by leveraging a Structure-Recognition-Relation (SRR) triplet paradigm. This design simplifies what would otherwise be a complex multi-tool pipeline and avoids the inefficiencies of processing full pages with giant end-to-end models. In SRR, document parsing is abstracted into three fundamental questions - ``Where is it?'' (structure), ``What is it?'' (recognition), and ``How is it organized?'' (relation) - corresponding to structure detection, content recognition, and relation prediction. To support this paradigm, we present MonkeyDoc, a comprehensive dataset with 4.5 million bilingual instances spanning over ten document types, which addresses the limitations of existing datasets that often focus on a single task, language, or document type. Leveraging the SRR paradigm and MonkeyDoc, we trained a 3B-parameter document foundation model. We further identify parameter redundancy in this model and propose contiguous parameter degradation (CPD), enabling the construction of models from 0.6B to 1.2B parameters that run faster with acceptable performance drop. MonkeyOCR achieves state-of-the-art performance, surpassing previous open-source and closed-source methods, including Gemini 2.5-Pro. Additionally, the model can be efficiently deployed for inference on a single RTX 3090 GPU. Code and models will be released at https://github.com/Yuliang-Liu/MonkeyOCR.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05218
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
Li, Zhang
Liu, Yuliang
Liu, Qiang
Ma, Zhiyin
Zhang, Ziyang
Zhang, Shuo
Yang, Biao
Guo, Zidun
Zhang, Jiarui
Wang, Xinyu
Bai, Xiang
Computer Vision and Pattern Recognition
We introduce MonkeyOCR, a document parsing model that advances the state of the art by leveraging a Structure-Recognition-Relation (SRR) triplet paradigm. This design simplifies what would otherwise be a complex multi-tool pipeline and avoids the inefficiencies of processing full pages with giant end-to-end models. In SRR, document parsing is abstracted into three fundamental questions - ``Where is it?'' (structure), ``What is it?'' (recognition), and ``How is it organized?'' (relation) - corresponding to structure detection, content recognition, and relation prediction. To support this paradigm, we present MonkeyDoc, a comprehensive dataset with 4.5 million bilingual instances spanning over ten document types, which addresses the limitations of existing datasets that often focus on a single task, language, or document type. Leveraging the SRR paradigm and MonkeyDoc, we trained a 3B-parameter document foundation model. We further identify parameter redundancy in this model and propose contiguous parameter degradation (CPD), enabling the construction of models from 0.6B to 1.2B parameters that run faster with acceptable performance drop. MonkeyOCR achieves state-of-the-art performance, surpassing previous open-source and closed-source methods, including Gemini 2.5-Pro. Additionally, the model can be efficiently deployed for inference on a single RTX 3090 GPU. Code and models will be released at https://github.com/Yuliang-Liu/MonkeyOCR.
title MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05218