DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Qintong, Zhang, Junyuan, Ren, Zhifei, Ouyang, Linke, Wen, Zichen, Niu, Junbo, Qu, Yuan, Wang, Bin, Chow, Ka-Ho, He, Conghui, Zhang, Wentao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918244348067840
author Zhang, Qintong
Zhang, Junyuan
Ren, Zhifei
Ouyang, Linke
Wen, Zichen
Niu, Junbo
Qu, Yuan
Wang, Bin
Chow, Ka-Ho
He, Conghui
Zhang, Wentao
author_facet Zhang, Qintong
Zhang, Junyuan
Ren, Zhifei
Ouyang, Linke
Wen, Zichen
Niu, Junbo
Qu, Yuan
Wang, Bin
Chow, Ka-Ho
He, Conghui
Zhang, Wentao
contents Document parsing aims to transform unstructured PDF images into semi-structured data, facilitating the digitization and utilization of information in diverse domains. While vision language models (VLMs) have significantly advanced this task, achieving reliable, high-quality parsing in real-world scenarios remains challenging. Common practice often selects the top-performing model on standard benchmarks. However, these benchmarks may carry dataset-specific biases, leading to inconsistent model rankings and limited correlation with real-world performance. Moreover, benchmark metrics typically provide only overall scores, which can obscure distinct error patterns in output. This raises a key challenge: how can we reliably and comprehensively assess document parsing quality in the wild? We address this problem with DOCR-Inspector, which formalizes document parsing assessment as fine-grained error detection and analysis. Leveraging VLM-as-a-Judge, DOCR-Inspector analyzes a document image and its parsed output, identifies all errors, assigns them to one of 28 predefined types, and produces a comprehensive quality assessment. To enable this capability, we construct DOCRcase-200K for training and propose the Chain-of-Checklist reasoning paradigm to enable the hierarchical structure of parsing quality assessment. For empirical validation, we introduce DOCRcaseBench, a set of 882 real-world document parsing cases with manual annotations. On this benchmark, DOCR-Inspector-7B outperforms commercial models like Gemini 2.5 Pro, as well as leading open-source models. Further experiments demonstrate that its quality assessments provide valuable guidance for parsing results refinement, making DOCR-Inspector both a practical evaluator and a driver for advancing document parsing systems at scale. Model and code are released at: https://github.com/ZZZZZQT/DOCR-Inspector.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
Zhang, Qintong
Zhang, Junyuan
Ren, Zhifei
Ouyang, Linke
Wen, Zichen
Niu, Junbo
Qu, Yuan
Wang, Bin
Chow, Ka-Ho
He, Conghui
Zhang, Wentao
Computer Vision and Pattern Recognition
Document parsing aims to transform unstructured PDF images into semi-structured data, facilitating the digitization and utilization of information in diverse domains. While vision language models (VLMs) have significantly advanced this task, achieving reliable, high-quality parsing in real-world scenarios remains challenging. Common practice often selects the top-performing model on standard benchmarks. However, these benchmarks may carry dataset-specific biases, leading to inconsistent model rankings and limited correlation with real-world performance. Moreover, benchmark metrics typically provide only overall scores, which can obscure distinct error patterns in output. This raises a key challenge: how can we reliably and comprehensively assess document parsing quality in the wild? We address this problem with DOCR-Inspector, which formalizes document parsing assessment as fine-grained error detection and analysis. Leveraging VLM-as-a-Judge, DOCR-Inspector analyzes a document image and its parsed output, identifies all errors, assigns them to one of 28 predefined types, and produces a comprehensive quality assessment. To enable this capability, we construct DOCRcase-200K for training and propose the Chain-of-Checklist reasoning paradigm to enable the hierarchical structure of parsing quality assessment. For empirical validation, we introduce DOCRcaseBench, a set of 882 real-world document parsing cases with manual annotations. On this benchmark, DOCR-Inspector-7B outperforms commercial models like Gemini 2.5 Pro, as well as leading open-source models. Further experiments demonstrate that its quality assessments provide valuable guidance for parsing results refinement, making DOCR-Inspector both a practical evaluator and a driver for advancing document parsing systems at scale. Model and code are released at: https://github.com/ZZZZZQT/DOCR-Inspector.
title DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.10619