MinerU: An Open-Source Solution for Precise Document Content Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912048489693184 |
|---|---|
| author | Wang, Bin Xu, Chao Zhao, Xiaomeng Ouyang, Linke Wu, Fan Zhao, Zhiyuan Xu, Rui Liu, Kaiwen Qu, Yuan Shang, Fukai Zhang, Bo Wei, Liqun Sui, Zhihao Li, Wei Shi, Botian Qiao, Yu Lin, Dahua He, Conghui |
| author_facet | Wang, Bin Xu, Chao Zhao, Xiaomeng Ouyang, Linke Wu, Fan Zhao, Zhiyuan Xu, Rui Liu, Kaiwen Qu, Yuan Shang, Fukai Zhang, Bo Wei, Liqun Sui, Zhihao Li, Wei Shi, Botian Qiao, Yu Lin, Dahua He, Conghui |
| contents | Document content analysis has been a crucial research area in computer vision. Despite significant advancements in methods such as OCR, layout detection, and formula recognition, existing open-source solutions struggle to consistently deliver high-quality content extraction due to the diversity in document types and content. To address these challenges, we present MinerU, an open-source solution for high-precision document content extraction. MinerU leverages the sophisticated PDF-Extract-Kit models to extract content from diverse documents effectively and employs finely-tuned preprocessing and postprocessing rules to ensure the accuracy of the final results. Experimental results demonstrate that MinerU consistently achieves high performance across various document types, significantly enhancing the quality and consistency of content extraction. The MinerU open-source project is available at https://github.com/opendatalab/MinerU. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_18839 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MinerU: An Open-Source Solution for Precise Document Content Extraction Wang, Bin Xu, Chao Zhao, Xiaomeng Ouyang, Linke Wu, Fan Zhao, Zhiyuan Xu, Rui Liu, Kaiwen Qu, Yuan Shang, Fukai Zhang, Bo Wei, Liqun Sui, Zhihao Li, Wei Shi, Botian Qiao, Yu Lin, Dahua He, Conghui Computer Vision and Pattern Recognition Document content analysis has been a crucial research area in computer vision. Despite significant advancements in methods such as OCR, layout detection, and formula recognition, existing open-source solutions struggle to consistently deliver high-quality content extraction due to the diversity in document types and content. To address these challenges, we present MinerU, an open-source solution for high-precision document content extraction. MinerU leverages the sophisticated PDF-Extract-Kit models to extract content from diverse documents effectively and employs finely-tuned preprocessing and postprocessing rules to ensure the accuracy of the final results. Experimental results demonstrate that MinerU consistently achieves high performance across various document types, significantly enhancing the quality and consistency of content extraction. The MinerU open-source project is available at https://github.com/opendatalab/MinerU. |
| title | MinerU: An Open-Source Solution for Precise Document Content Extraction |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2409.18839 |