DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qintong, Zhang, Junyuan, Ren, Zhifei, Ouyang, Linke, Wen, Zichen, Niu, Junbo, Qu, Yuan, Wang, Bin, Chow, Ka-Ho, He, Conghui, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
by: Zhang, Junyuan, et al.
Published: (2024)
by: Zhang, Junyuan, et al.
Published: (2024)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
by: Zhang, Qintong, et al.
Published: (2024)
by: Zhang, Qintong, et al.
Published: (2024)
TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition
by: Zhang, Junyuan, et al.
Published: (2025)
by: Zhang, Junyuan, et al.
Published: (2025)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
by: Dong, Hejun, et al.
Published: (2026)
by: Dong, Hejun, et al.
Published: (2026)
Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
by: Ouyang, Linke, et al.
Published: (2024)
by: Ouyang, Linke, et al.
Published: (2024)
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
by: Niu, Junbo, et al.
Published: (2025)
by: Niu, Junbo, et al.
Published: (2025)
MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
by: Wang, Bin, et al.
Published: (2026)
by: Wang, Bin, et al.
Published: (2026)
C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
by: Wang, Zhengren, et al.
Published: (2026)
by: Wang, Zhengren, et al.
Published: (2026)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
Confundo: Learning to Generate Robust Poison for Practical RAG Systems
by: Hu, Haoyang, et al.
Published: (2026)
by: Hu, Haoyang, et al.
Published: (2026)
Building Gradient Bridges: Label Leakage from Restricted Gradient Sharing in Federated Learning
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
by: Lai, Peichao, et al.
Published: (2025)
by: Lai, Peichao, et al.
Published: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
by: Kang, Hengrui, et al.
Published: (2025)
by: Kang, Hengrui, et al.
Published: (2025)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
by: Cui, Cheng, et al.
Published: (2026)
by: Cui, Cheng, et al.
Published: (2026)
Harmless Backdoor-based Client-side Watermarking in Federated Learning
by: Luo, Kaijing, et al.
Published: (2024)
by: Luo, Kaijing, et al.
Published: (2024)
Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
by: Zhao, Zhiyuan, et al.
Published: (2023)
by: Zhao, Zhiyuan, et al.
Published: (2023)
Fine-Grained VLM Fine-tuning via Latent Hierarchical Adapter Learning
by: Zhao, Yumiao, et al.
Published: (2025)
by: Zhao, Yumiao, et al.
Published: (2025)
DocDancer: Towards Agentic Document-Grounded Information Seeking
by: Zhang, Qintong, et al.
Published: (2026)
by: Zhang, Qintong, et al.
Published: (2026)
Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
Fine structure of rupture set for semilinear elliptic equation with singular nonlinearity
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Attention with Dependency Parsing Augmentation for Fine-Grained Attribution
by: Ding, Qiang, et al.
Published: (2024)
by: Ding, Qiang, et al.
Published: (2024)
HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
by: Wu, Junhao, et al.
Published: (2025)
by: Wu, Junhao, et al.
Published: (2025)
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
by: Cui, Cheng, et al.
Published: (2026)
by: Cui, Cheng, et al.
Published: (2026)
Linear inviscid damping in the presence of an embedding eigenvalue
by: Ren, Siqi, et al.
Published: (2024)
by: Ren, Siqi, et al.
Published: (2024)
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
by: Ma, Yiming, et al.
Published: (2026)
by: Ma, Yiming, et al.
Published: (2026)
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
by: Xu, Bangrui, et al.
Published: (2026)
by: Xu, Bangrui, et al.
Published: (2026)
Graph Neural Networks with Coarse- and Fine-Grained Division for Mitigating Label Sparsity and Noise
by: Li, Shuangjie, et al.
Published: (2024)
by: Li, Shuangjie, et al.
Published: (2024)
Imperio: Language-Guided Backdoor Attacks for Arbitrary Model Control
by: Chow, Ka-Ho, et al.
Published: (2024)
by: Chow, Ka-Ho, et al.
Published: (2024)
ParseBench: A Document Parsing Benchmark for AI Agents
by: Zhang, Boyang, et al.
Published: (2026)
by: Zhang, Boyang, et al.
Published: (2026)
FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval
by: Li, Zixu, et al.
Published: (2025)
by: Li, Zixu, et al.
Published: (2025)
github.com/FusionInspector/FusionInspector/fusion_inspector
by: Brian Haas
Published: (2026)
by: Brian Haas
Published: (2026)
github.com/FusionInspector/FusionInspector/fusion_inspector
by: Brian Haas
Published: (2026)
by: Brian Haas
Published: (2026)
On the Efficiency of Privacy Attacks in Federated Learning
by: Tabassum, Nawrin, et al.
Published: (2024)
by: Tabassum, Nawrin, et al.
Published: (2024)
DocFusion: A Unified Framework for Document Parsing Tasks
by: Chai, Mingxu, et al.
Published: (2024)
by: Chai, Mingxu, et al.
Published: (2024)
Molecular Identifier Visual Prompt and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing
by: Song, Jiahe, et al.
Published: (2026)
by: Song, Jiahe, et al.
Published: (2026)
Similar Items
-
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
by: Zhang, Junyuan, et al.
Published: (2024) -
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
by: Zhang, Qintong, et al.
Published: (2024) -
TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition
by: Zhang, Junyuan, et al.
Published: (2025) -
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
by: Wen, Zichen, et al.
Published: (2025) -
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
by: Dong, Hejun, et al.
Published: (2026)