Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918534395723776 |
|---|---|
| author | Yang, Minglai Yu, Xinyan Velocity Li, Pengyuan Guo, Xinyu Qi, Zhenting Kim, Konwoo Ye, Longtian Luo, Xiaolong Bi, Jinhe Zhang, Henry Riaz, Haris Zhang, Xuan Xiao, Yunze Liu, Bangya Tang, Tom Zhao, Yunfei Lin, Qunshu Wang, Zihan Liu, Minghao Li, Michael Lingzhi Du, Yilun Thomason, Jesse Feris, Rogerio Pentland, Alex He, Zexue |
| author_facet | Yang, Minglai Yu, Xinyan Velocity Li, Pengyuan Guo, Xinyu Qi, Zhenting Kim, Konwoo Ye, Longtian Luo, Xiaolong Bi, Jinhe Zhang, Henry Riaz, Haris Zhang, Xuan Xiao, Yunze Liu, Bangya Tang, Tom Zhao, Yunfei Lin, Qunshu Wang, Zihan Liu, Minghao Li, Michael Lingzhi Du, Yilun Thomason, Jesse Feris, Rogerio Pentland, Alex He, Zexue |
| contents | Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are increasingly limited in coverage and difficulty: many focus on common document genres or uniformly sampled pages where modern parsers already perform strongly, while offering limited annotation for expert-domain structures such as chemical formula, music notation, complex tables, and cross-page layouts. We introduce Dr. DocBench, a difficulty-aware benchmark for expert-level document parsing. Built from a large-scale multilingual book corpus, Dr. DocBench spans 52 BISAC subject domains and selects challenging documents through parser-failure-based sampling, targeting cases where multiple state-of-the-art systems struggle. It contains 4,514 annotated pages from long documents averaging around 100 pages, with 65k high-quality page- and block-level annotations for layout, reading order, hierarchical relations, and domain-specific visual contents. Evaluations of pipeline-based parsers and general-purpose VLMs show that strong performance on existing benchmarks does not transfer to our expert-level document parsing. Our analysis reveals substantial failures across subjects, content types, and structural attributes, highlighting Dr. DocBench as a comprehensive testbed for diagnosing and advancing document intelligence. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2606_01393 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing Yang, Minglai Yu, Xinyan Velocity Li, Pengyuan Guo, Xinyu Qi, Zhenting Kim, Konwoo Ye, Longtian Luo, Xiaolong Bi, Jinhe Zhang, Henry Riaz, Haris Zhang, Xuan Xiao, Yunze Liu, Bangya Tang, Tom Zhao, Yunfei Lin, Qunshu Wang, Zihan Liu, Minghao Li, Michael Lingzhi Du, Yilun Thomason, Jesse Feris, Rogerio Pentland, Alex He, Zexue Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are increasingly limited in coverage and difficulty: many focus on common document genres or uniformly sampled pages where modern parsers already perform strongly, while offering limited annotation for expert-domain structures such as chemical formula, music notation, complex tables, and cross-page layouts. We introduce Dr. DocBench, a difficulty-aware benchmark for expert-level document parsing. Built from a large-scale multilingual book corpus, Dr. DocBench spans 52 BISAC subject domains and selects challenging documents through parser-failure-based sampling, targeting cases where multiple state-of-the-art systems struggle. It contains 4,514 annotated pages from long documents averaging around 100 pages, with 65k high-quality page- and block-level annotations for layout, reading order, hierarchical relations, and domain-specific visual contents. Evaluations of pipeline-based parsers and general-purpose VLMs show that strong performance on existing benchmarks does not transfer to our expert-level document parsing. Our analysis reveals substantial failures across subjects, content types, and structural attributes, highlighting Dr. DocBench as a comprehensive testbed for diagnosing and advancing document intelligence. |
| title | Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2606.01393 |