MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiarui, Liu, Yuliang, Wu, Zijun, Pang, Guosheng, Ye, Zhili, Zhong, Yupei, Ma, Junteng, Wei, Tao, Xu, Haiyang, Chen, Weikai, Wang, Zeen, Ji, Qiangjun, Zhou, Fanxi, Zhang, Qi, Hu, Yuanrui, Liu, Jiahao, Li, Zhang, Zhang, Ziyang, Liu, Qiang, Bai, Xiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911267722100736
author Zhang, Jiarui
Liu, Yuliang
Wu, Zijun
Pang, Guosheng
Ye, Zhili
Zhong, Yupei
Ma, Junteng
Wei, Tao
Xu, Haiyang
Chen, Weikai
Wang, Zeen
Ji, Qiangjun
Zhou, Fanxi
Zhang, Qi
Hu, Yuanrui
Liu, Jiahao
Li, Zhang
Zhang, Ziyang
Liu, Qiang
Bai, Xiang
author_facet Zhang, Jiarui
Liu, Yuliang
Wu, Zijun
Pang, Guosheng
Ye, Zhili
Zhong, Yupei
Ma, Junteng
Wei, Tao
Xu, Haiyang
Chen, Weikai
Wang, Zeen
Ji, Qiangjun
Zhou, Fanxi
Zhang, Qi
Hu, Yuanrui
Liu, Jiahao
Li, Zhang
Zhang, Ziyang
Liu, Qiang
Bai, Xiang
contents Document parsing is a core task in document intelligence, supporting applications such as information extraction, retrieval-augmented generation, and automated document analysis. However, real-world documents often feature complex layouts with multi-level tables, embedded images or formulas, and cross-page structures, which remain challenging for existing OCR systems. We introduce MonkeyOCR v1.5, a unified vision-language framework that enhances both layout understanding and content recognition through a two-stage pipeline. The first stage employs a large multimodal model to jointly predict layout and reading order, leveraging visual information to ensure sequential consistency. The second stage performs localized recognition of text, formulas, and tables within detected regions, maintaining high visual fidelity while reducing error propagation. To address complex table structures, we propose a visual consistency-based reinforcement learning scheme that evaluates recognition quality via render-and-compare alignment, improving structural accuracy without manual annotations. Additionally, two specialized modules, Image-Decoupled Table Parsing and Type-Guided Table Merging, are introduced to enable reliable parsing of tables containing embedded images and reconstruction of tables crossing pages or columns. Comprehensive experiments on OmniDocBench v1.5 demonstrate that MonkeyOCR v1.5 achieves state-of-the-art performance, outperforming PPOCR-VL and MinerU 2.5 while showing exceptional robustness in visually complex document scenarios. A trial link can be found at https://github.com/Yuliang-Liu/MonkeyOCR .
format Preprint
id arxiv_https___arxiv_org_abs_2511_10390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
Zhang, Jiarui
Liu, Yuliang
Wu, Zijun
Pang, Guosheng
Ye, Zhili
Zhong, Yupei
Ma, Junteng
Wei, Tao
Xu, Haiyang
Chen, Weikai
Wang, Zeen
Ji, Qiangjun
Zhou, Fanxi
Zhang, Qi
Hu, Yuanrui
Liu, Jiahao
Li, Zhang
Zhang, Ziyang
Liu, Qiang
Bai, Xiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Document parsing is a core task in document intelligence, supporting applications such as information extraction, retrieval-augmented generation, and automated document analysis. However, real-world documents often feature complex layouts with multi-level tables, embedded images or formulas, and cross-page structures, which remain challenging for existing OCR systems. We introduce MonkeyOCR v1.5, a unified vision-language framework that enhances both layout understanding and content recognition through a two-stage pipeline. The first stage employs a large multimodal model to jointly predict layout and reading order, leveraging visual information to ensure sequential consistency. The second stage performs localized recognition of text, formulas, and tables within detected regions, maintaining high visual fidelity while reducing error propagation. To address complex table structures, we propose a visual consistency-based reinforcement learning scheme that evaluates recognition quality via render-and-compare alignment, improving structural accuracy without manual annotations. Additionally, two specialized modules, Image-Decoupled Table Parsing and Type-Guided Table Merging, are introduced to enable reliable parsing of tables containing embedded images and reconstruction of tables crossing pages or columns. Comprehensive experiments on OmniDocBench v1.5 demonstrate that MonkeyOCR v1.5 achieves state-of-the-art performance, outperforming PPOCR-VL and MinerU 2.5 while showing exceptional robustness in visually complex document scenarios. A trial link can be found at https://github.com/Yuliang-Liu/MonkeyOCR .
title MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.10390