How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhiheng, Ma, Zongyang, Chen, Jiaxian, Zhang, Jianing, Su, Zhaolong, Zhang, Yutong, Yu, Zhiyin, Liu, Ruiqi, Lv, Xiaolei, Li, Bo, Gao, Jun, Zhang, Ziqi, Yuan, Chunfeng, Li, Bing, Hu, Weiming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914543391735808
author Li, Zhiheng
Ma, Zongyang
Chen, Jiaxian
Zhang, Jianing
Su, Zhaolong
Zhang, Yutong
Yu, Zhiyin
Liu, Ruiqi
Lv, Xiaolei
Li, Bo
Gao, Jun
Zhang, Ziqi
Yuan, Chunfeng
Li, Bing
Hu, Weiming
author_facet Li, Zhiheng
Ma, Zongyang
Chen, Jiaxian
Zhang, Jianing
Su, Zhaolong
Zhang, Yutong
Yu, Zhiyin
Liu, Ruiqi
Lv, Xiaolei
Li, Bo
Gao, Jun
Zhang, Ziqi
Yuan, Chunfeng
Li, Bing
Hu, Weiming
contents The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose top scores have saturated above 90%. Athree-stage audit pipeline we run on OmniDocBench screens its 21,353evaluator-scored blocks and confirms 2,580 errors (12.08%); combined with overa year of public availability, both annotation quality and contamination riskcall its rankings into question. To address these issues, we presentPureDocBench, a programmatically generated, source-traceable benchmark thatrenders document images from HTML/CSS and produces verifiable annotations fromthe same source, covering 10 domains, 66 subcategories, and 1,475 pages, eachin three versions: clean, digitally degraded, and real-degraded (4,425 imagestotal). Evaluating 40 models spanning pipeline specialists, end-to-endspecialists, and general-purpose VLMs, we find: (i) document parsing is farfrom solved: the best model scores only ~74 out of 100, with a 44.6-point gapbetween the strongest and weakest models; (ii) specialist parsers with <=4Bparameters rival or surpass general VLMs that are 5-100x larger, yet formularecognition remains a shared bottleneck where no model exceeds 67% whenaveraging the formula metric across all three tracks; (iii) general VLMs loseonly 0.99/8.52 Overall points under digital/real degradation versus 4.90/14.21for pipeline specialists, producing ranking reversals that make clean-onlyevaluation misleading for deployment. All data, code, and artifacts arepublicly released.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07492
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings
Li, Zhiheng
Ma, Zongyang
Chen, Jiaxian
Zhang, Jianing
Su, Zhaolong
Zhang, Yutong
Yu, Zhiyin
Liu, Ruiqi
Lv, Xiaolei
Li, Bo
Gao, Jun
Zhang, Ziqi
Yuan, Chunfeng
Li, Bing
Hu, Weiming
Computer Vision and Pattern Recognition
The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose top scores have saturated above 90%. Athree-stage audit pipeline we run on OmniDocBench screens its 21,353evaluator-scored blocks and confirms 2,580 errors (12.08%); combined with overa year of public availability, both annotation quality and contamination riskcall its rankings into question. To address these issues, we presentPureDocBench, a programmatically generated, source-traceable benchmark thatrenders document images from HTML/CSS and produces verifiable annotations fromthe same source, covering 10 domains, 66 subcategories, and 1,475 pages, eachin three versions: clean, digitally degraded, and real-degraded (4,425 imagestotal). Evaluating 40 models spanning pipeline specialists, end-to-endspecialists, and general-purpose VLMs, we find: (i) document parsing is farfrom solved: the best model scores only ~74 out of 100, with a 44.6-point gapbetween the strongest and weakest models; (ii) specialist parsers with <=4Bparameters rival or surpass general VLMs that are 5-100x larger, yet formularecognition remains a shared bottleneck where no model exceeds 67% whenaveraging the formula metric across all three tracks; (iii) general VLMs loseonly 0.99/8.52 Overall points under digital/real degradation versus 4.90/14.21for pipeline specialists, producing ranking reversals that make clean-onlyevaluation misleading for deployment. All data, code, and artifacts arepublicly released.
title How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.07492