Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Changda, Gao, Ziyue, Wang, Xueqing, Gao, Tingquan, Cui, Cheng, Tang, Jing, Liu, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings
von: Li, Zhiheng, et al.
Veröffentlicht: (2026)
von: Li, Zhiheng, et al.
Veröffentlicht: (2026)
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
von: Zhou, Bangbang, et al.
Veröffentlicht: (2026)
von: Zhou, Bangbang, et al.
Veröffentlicht: (2026)
PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
ParseBench: A Document Parsing Benchmark for AI Agents
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
von: Wang, An-Lan, et al.
Veröffentlicht: (2025)
von: Wang, An-Lan, et al.
Veröffentlicht: (2025)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
DocFusion: A Unified Framework for Document Parsing Tasks
von: Chai, Mingxu, et al.
Veröffentlicht: (2024)
von: Chai, Mingxu, et al.
Veröffentlicht: (2024)
Logics-Parsing-Omni Technical Report
von: An, Xin, et al.
Veröffentlicht: (2026)
von: An, Xin, et al.
Veröffentlicht: (2026)
PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction
von: Sun, Ting, et al.
Veröffentlicht: (2025)
von: Sun, Ting, et al.
Veröffentlicht: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
RenoBench: A Citation Parsing Benchmark
von: Sarin, Parth, et al.
Veröffentlicht: (2026)
von: Sarin, Parth, et al.
Veröffentlicht: (2026)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagery
von: Tang, Jiadong, et al.
Veröffentlicht: (2025)
von: Tang, Jiadong, et al.
Veröffentlicht: (2025)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
von: Li, Zhang, et al.
Veröffentlicht: (2026)
von: Li, Zhang, et al.
Veröffentlicht: (2026)
DocParseNet: Advanced Semantic Segmentation and OCR Embeddings for Efficient Scanned Document Annotation
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
von: Chen, Ketong, et al.
Veröffentlicht: (2025)
von: Chen, Ketong, et al.
Veröffentlicht: (2025)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data
von: Wang, Xuwu, et al.
Veröffentlicht: (2024)
von: Wang, Xuwu, et al.
Veröffentlicht: (2024)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
von: Gao, Zeyu, et al.
Veröffentlicht: (2025)
von: Gao, Zeyu, et al.
Veröffentlicht: (2025)
PaddleOCR 3.0 Technical Report
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
von: Gao, Hong, et al.
Veröffentlicht: (2025)
von: Gao, Hong, et al.
Veröffentlicht: (2025)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
von: Deng, Chao, et al.
Veröffentlicht: (2024)
von: Deng, Chao, et al.
Veröffentlicht: (2024)
Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
von: Yang, Yujia, et al.
Veröffentlicht: (2026)
von: Yang, Yujia, et al.
Veröffentlicht: (2026)
CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance
von: Pai, Kunal, et al.
Veröffentlicht: (2025)
von: Pai, Kunal, et al.
Veröffentlicht: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
von: Wang, Pengyu, et al.
Veröffentlicht: (2026)
von: Wang, Pengyu, et al.
Veröffentlicht: (2026)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
von: Zhang, Linhao, et al.
Veröffentlicht: (2025)
von: Zhang, Linhao, et al.
Veröffentlicht: (2025)
FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation
von: Zhu, Jiayong, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayong, et al.
Veröffentlicht: (2026)
DocTabQA: Answering Questions from Long Documents Using Tables
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
von: Wang, Haochen, et al.
Veröffentlicht: (2024)
MolRecBench-Wild: A Real-World Benchmark for Optical Chemical Structure Recognition
von: Yang, Haote, et al.
Veröffentlicht: (2026)
von: Yang, Haote, et al.
Veröffentlicht: (2026)
Manipulate as Human: Learning Task-oriented Manipulation Skills by Adversarial Motion Priors
von: Ma, Ziqi, et al.
Veröffentlicht: (2025)
von: Ma, Ziqi, et al.
Veröffentlicht: (2025)
MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
von: Cui, Junbo, et al.
Veröffentlicht: (2026)
von: Cui, Junbo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
von: Ouyang, Linke, et al.
Veröffentlicht: (2024) -
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
von: Cui, Cheng, et al.
Veröffentlicht: (2026) -
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
von: Cui, Cheng, et al.
Veröffentlicht: (2026) -
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
von: Yang, Minglai, et al.
Veröffentlicht: (2026) -
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
von: Cui, Cheng, et al.
Veröffentlicht: (2025)