Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Minglai, Yu, Xinyan Velocity, Li, Pengyuan, Guo, Xinyu, Qi, Zhenting, Kim, Konwoo, Ye, Longtian, Luo, Xiaolong, Bi, Jinhe, Zhang, Henry, Riaz, Haris, Zhang, Xuan, Xiao, Yunze, Liu, Bangya, Tang, Tom, Zhao, Yunfei, Lin, Qunshu, Wang, Zihan, Liu, Minghao, Li, Michael Lingzhi, Du, Yilun, Thomason, Jesse, Feris, Rogerio, Pentland, Alex, He, Zexue
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918534395723776
author Yang, Minglai
Yu, Xinyan Velocity
Li, Pengyuan
Guo, Xinyu
Qi, Zhenting
Kim, Konwoo
Ye, Longtian
Luo, Xiaolong
Bi, Jinhe
Zhang, Henry
Riaz, Haris
Zhang, Xuan
Xiao, Yunze
Liu, Bangya
Tang, Tom
Zhao, Yunfei
Lin, Qunshu
Wang, Zihan
Liu, Minghao
Li, Michael Lingzhi
Du, Yilun
Thomason, Jesse
Feris, Rogerio
Pentland, Alex
He, Zexue
author_facet Yang, Minglai
Yu, Xinyan Velocity
Li, Pengyuan
Guo, Xinyu
Qi, Zhenting
Kim, Konwoo
Ye, Longtian
Luo, Xiaolong
Bi, Jinhe
Zhang, Henry
Riaz, Haris
Zhang, Xuan
Xiao, Yunze
Liu, Bangya
Tang, Tom
Zhao, Yunfei
Lin, Qunshu
Wang, Zihan
Liu, Minghao
Li, Michael Lingzhi
Du, Yilun
Thomason, Jesse
Feris, Rogerio
Pentland, Alex
He, Zexue
contents Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are increasingly limited in coverage and difficulty: many focus on common document genres or uniformly sampled pages where modern parsers already perform strongly, while offering limited annotation for expert-domain structures such as chemical formula, music notation, complex tables, and cross-page layouts. We introduce Dr. DocBench, a difficulty-aware benchmark for expert-level document parsing. Built from a large-scale multilingual book corpus, Dr. DocBench spans 52 BISAC subject domains and selects challenging documents through parser-failure-based sampling, targeting cases where multiple state-of-the-art systems struggle. It contains 4,514 annotated pages from long documents averaging around 100 pages, with 65k high-quality page- and block-level annotations for layout, reading order, hierarchical relations, and domain-specific visual contents. Evaluations of pipeline-based parsers and general-purpose VLMs show that strong performance on existing benchmarks does not transfer to our expert-level document parsing. Our analysis reveals substantial failures across subjects, content types, and structural attributes, highlighting Dr. DocBench as a comprehensive testbed for diagnosing and advancing document intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01393
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
Yang, Minglai
Yu, Xinyan Velocity
Li, Pengyuan
Guo, Xinyu
Qi, Zhenting
Kim, Konwoo
Ye, Longtian
Luo, Xiaolong
Bi, Jinhe
Zhang, Henry
Riaz, Haris
Zhang, Xuan
Xiao, Yunze
Liu, Bangya
Tang, Tom
Zhao, Yunfei
Lin, Qunshu
Wang, Zihan
Liu, Minghao
Li, Michael Lingzhi
Du, Yilun
Thomason, Jesse
Feris, Rogerio
Pentland, Alex
He, Zexue
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are increasingly limited in coverage and difficulty: many focus on common document genres or uniformly sampled pages where modern parsers already perform strongly, while offering limited annotation for expert-domain structures such as chemical formula, music notation, complex tables, and cross-page layouts. We introduce Dr. DocBench, a difficulty-aware benchmark for expert-level document parsing. Built from a large-scale multilingual book corpus, Dr. DocBench spans 52 BISAC subject domains and selects challenging documents through parser-failure-based sampling, targeting cases where multiple state-of-the-art systems struggle. It contains 4,514 annotated pages from long documents averaging around 100 pages, with 65k high-quality page- and block-level annotations for layout, reading order, hierarchical relations, and domain-specific visual contents. Evaluations of pipeline-based parsers and general-purpose VLMs show that strong performance on existing benchmarks does not transfer to our expert-level document parsing. Our analysis reveals substantial failures across subjects, content types, and structural attributes, highlighting Dr. DocBench as a comprehensive testbed for diagnosing and advancing document intelligence.
title Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2606.01393