Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Minglai, Yu, Xinyan Velocity, Li, Pengyuan, Guo, Xinyu, Qi, Zhenting, Kim, Konwoo, Ye, Longtian, Luo, Xiaolong, Bi, Jinhe, Zhang, Henry, Riaz, Haris, Zhang, Xuan, Xiao, Yunze, Liu, Bangya, Tang, Tom, Zhao, Yunfei, Lin, Qunshu, Wang, Zihan, Liu, Minghao, Li, Michael Lingzhi, Du, Yilun, Thomason, Jesse, Feris, Rogerio, Pentland, Alex, He, Zexue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
by: Zhu, Xiaoyuan, et al.
Published: (2026)
by: Zhu, Xiaoyuan, et al.
Published: (2026)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
by: Ouyang, Linke, et al.
Published: (2024)
by: Ouyang, Linke, et al.
Published: (2024)
Composition-Grounded Data Synthesis for Visual Reasoning
by: Gu, Xinyi, et al.
Published: (2025)
by: Gu, Xinyi, et al.
Published: (2025)
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
by: Liu, Bangya, et al.
Published: (2024)
by: Liu, Bangya, et al.
Published: (2024)
DocFusion: A Unified Framework for Document Parsing Tasks
by: Chai, Mingxu, et al.
Published: (2024)
by: Chai, Mingxu, et al.
Published: (2024)
Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research
by: Dong, Kuicai, et al.
Published: (2025)
by: Dong, Kuicai, et al.
Published: (2025)
CAMELoT: Towards Large Language Models with Training-Free Consolidated Associative Memory
by: He, Zexue, et al.
Published: (2024)
by: He, Zexue, et al.
Published: (2024)
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
by: Wang, Wenjie, et al.
Published: (2026)
by: Wang, Wenjie, et al.
Published: (2026)
Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation
by: Riaz, Haris, et al.
Published: (2025)
by: Riaz, Haris, et al.
Published: (2025)
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
by: Kondic, Jovana, et al.
Published: (2026)
by: Kondic, Jovana, et al.
Published: (2026)
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
by: Zhou, Changda, et al.
Published: (2026)
by: Zhou, Changda, et al.
Published: (2026)
Images Inpainting Quality Evaluation Using Structural Features and Visual Saliency
by: Shuang Ma, et al.
Published: (2024)
by: Shuang Ma, et al.
Published: (2024)
AlignSAE: Concept-Aligned Sparse Autoencoders
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
DocParseNet: Advanced Semantic Segmentation and OCR Embeddings for Efficient Scanned Document Annotation
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
Pre-training under infinite compute
by: Kim, Konwoo, et al.
Published: (2025)
by: Kim, Konwoo, et al.
Published: (2025)
ELLEN: Extremely Lightly Supervised Learning For Efficient Named Entity Recognition
by: Riaz, Haris, et al.
Published: (2024)
by: Riaz, Haris, et al.
Published: (2024)
MV-S2V: Multi-View Subject-Consistent Video Generation
by: Song, Ziyang, et al.
Published: (2026)
by: Song, Ziyang, et al.
Published: (2026)
Syntomification and crystalline local systems
by: Pentland, Dylan
Published: (2025)
by: Pentland, Dylan
Published: (2025)
White Nanozymes with Enhanced Alkaline Phosphatase‐Mimicking Activity and Selective Inhibition Effect: Enzyme‐Free Colorimetric Test Strip of Pesticide
by: Xinyan Guo, et al.
Published: (2025)
by: Xinyan Guo, et al.
Published: (2025)
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
by: Liu, Xuan, et al.
Published: (2026)
by: Liu, Xuan, et al.
Published: (2026)
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
by: Chen, Weidong, et al.
Published: (2026)
by: Chen, Weidong, et al.
Published: (2026)
PhysLab: A Benchmark Dataset for Multi-Granularity Visual Parsing of Physics Experiments
by: Zou, Minghao, et al.
Published: (2025)
by: Zou, Minghao, et al.
Published: (2025)
Difficult Times, Difficult People
by: Johnson, Doug
Published: (2004)
by: Johnson, Doug
Published: (2004)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
Human Drive System Dynamics (HDSD): A Metatheoretical Framework
by: Jinhe, Hao
Published: (2025)
by: Jinhe, Hao
Published: (2025)
A note on $μ$-stabilizers in ACVF
by: Ye, Jinhe
Published: (2019)
by: Ye, Jinhe
Published: (2019)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
by: Kondic, Jovana, et al.
Published: (2025)
by: Kondic, Jovana, et al.
Published: (2025)
Writeaerobics: 40 Workshop Exercises To Improve Your Writing Teaching. Bill Harp Professional Teachers Library.
by: Thomason, Tommy
Published: (2003)
by: Thomason, Tommy
Published: (2003)
Writer to Writer: How To Conference Young Authors. The Bill Harp Professional Teachers Library Series.
by: Thomason, Tommy
Published: (1998)
by: Thomason, Tommy
Published: (1998)
Comparing Librarian and Student Assessments of Community College Library Services in Texas
by: Thomason, Nevada
Published: (1976)
by: Thomason, Nevada
Published: (1976)
Revitalization of Library Service.
by: Thomason, Jean
Published: (1993)
by: Thomason, Jean
Published: (1993)
Microcomputers and Automation in the School Library Media Center.
by: Thomason, Nevada
Published: (1982)
by: Thomason, Nevada
Published: (1982)
M+: Extending MemoryLLM with Scalable Long-Term Memory
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Objaverse++: Curated 3D Object Dataset with Quality Annotations
by: Lin, Chendi, et al.
Published: (2025)
by: Lin, Chendi, et al.
Published: (2025)
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
Quantum Wiretap Channel Coding Assisted by Noisy Correlation
by: Cai, Minglai, et al.
Published: (2024)
by: Cai, Minglai, et al.
Published: (2024)
Quantum Byzantine Multiple Access Channels
by: Cai, Minglai, et al.
Published: (2025)
by: Cai, Minglai, et al.
Published: (2025)
Word2VecGD: Neural Graph Drawing with Cosine-Stress Optimization
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
Similar Items
-
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
by: Zhu, Xiaoyuan, et al.
Published: (2026) -
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
by: Ouyang, Linke, et al.
Published: (2024) -
Composition-Grounded Data Synthesis for Visual Reasoning
by: Gu, Xinyi, et al.
Published: (2025) -
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
by: Liu, Bangya, et al.
Published: (2024) -
DocFusion: A Unified Framework for Document Parsing Tasks
by: Chai, Mingxu, et al.
Published: (2024)