MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Bangrui, Miao, Ziyang, Zhou, Xuanhe, Lin, Yiming, Tang, Zirui, Zhao, Xiaomeng, Wu, Fan, Tan, Cheng, Wang, Bin, He, Conghui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
por: Wang, Bin, et al.
Publicado: (2026)
por: Wang, Bin, et al.
Publicado: (2026)
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
por: Niu, Junbo, et al.
Publicado: (2025)
por: Niu, Junbo, et al.
Publicado: (2025)
MinerU: An Open-Source Solution for Precise Document Content Extraction
por: Wang, Bin, et al.
Publicado: (2024)
por: Wang, Bin, et al.
Publicado: (2024)
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
por: Dong, Hejun, et al.
Publicado: (2026)
por: Dong, Hejun, et al.
Publicado: (2026)
MoDora: Tree-Based Semi-Structured Document Analysis System
por: Xu, Bangrui, et al.
Publicado: (2026)
por: Xu, Bangrui, et al.
Publicado: (2026)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
por: Zhang, Qintong, et al.
Publicado: (2024)
por: Zhang, Qintong, et al.
Publicado: (2024)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
por: Ouyang, Linke, et al.
Publicado: (2024)
por: Ouyang, Linke, et al.
Publicado: (2024)
DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
por: Zhang, Qintong, et al.
Publicado: (2025)
por: Zhang, Qintong, et al.
Publicado: (2025)
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
por: Tian, Juanxi, et al.
Publicado: (2025)
por: Tian, Juanxi, et al.
Publicado: (2025)
Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
por: Feng, Hao, et al.
Publicado: (2026)
por: Feng, Hao, et al.
Publicado: (2026)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
por: Wang, Zhengren, et al.
Publicado: (2026)
por: Wang, Zhengren, et al.
Publicado: (2026)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
por: Li, Zhang, et al.
Publicado: (2025)
por: Li, Zhang, et al.
Publicado: (2025)
Logics-Parsing Technical Report
por: Chen, Xiangyang, et al.
Publicado: (2025)
por: Chen, Xiangyang, et al.
Publicado: (2025)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
por: Li, Huilai, et al.
Publicado: (2026)
por: Li, Huilai, et al.
Publicado: (2026)
Molecular Identifier Visual Prompt and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing
por: Song, Jiahe, et al.
Publicado: (2026)
por: Song, Jiahe, et al.
Publicado: (2026)
Document Author Classification Using Parsed Language Structure
por: Moon, Todd K, et al.
Publicado: (2024)
por: Moon, Todd K, et al.
Publicado: (2024)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
por: Wang, Wenjie, et al.
Publicado: (2026)
por: Wang, Wenjie, et al.
Publicado: (2026)
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
por: Feng, Hao, et al.
Publicado: (2025)
por: Feng, Hao, et al.
Publicado: (2025)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
por: Wang, Bin, et al.
Publicado: (2024)
por: Wang, Bin, et al.
Publicado: (2024)
ParseBench: A Document Parsing Benchmark for AI Agents
por: Zhang, Boyang, et al.
Publicado: (2026)
por: Zhang, Boyang, et al.
Publicado: (2026)
Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing
por: Fan, Jiahe, et al.
Publicado: (2026)
por: Fan, Jiahe, et al.
Publicado: (2026)
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
por: Cui, Cheng, et al.
Publicado: (2026)
por: Cui, Cheng, et al.
Publicado: (2026)
Geometry-aware Distance Measure for Diverse Hierarchical Structures in Hyperbolic Spaces
por: Li, Pengxiang, et al.
Publicado: (2025)
por: Li, Pengxiang, et al.
Publicado: (2025)
Parsing of Research Documents into XML Using Formal Grammars
por: Opeoluwa Iwashokun, et al.
Publicado: (2024)
por: Opeoluwa Iwashokun, et al.
Publicado: (2024)
Learning AND-OR Templates for Professional Photograph Parsing and Guidance
por: Jin, Xin, et al.
Publicado: (2024)
por: Jin, Xin, et al.
Publicado: (2024)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
por: Li, Zhang, et al.
Publicado: (2026)
por: Li, Zhang, et al.
Publicado: (2026)
AceParse: A Comprehensive Dataset with Diverse Structured Texts for Academic Literature Parsing
por: Ji, Huawei, et al.
Publicado: (2024)
por: Ji, Huawei, et al.
Publicado: (2024)
Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
por: Zhou, Wei, et al.
Publicado: (2026)
por: Zhou, Wei, et al.
Publicado: (2026)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
por: Wang, Baode, et al.
Publicado: (2025)
por: Wang, Baode, et al.
Publicado: (2025)
Lexicalized Constituency Parsing for Middle Dutch: Low-resource Training and Cross-Domain Generalization
por: Liang, Yiming, et al.
Publicado: (2026)
por: Liang, Yiming, et al.
Publicado: (2026)
Efficient Document Parsing via Parallel Token Prediction
por: Li, Lei, et al.
Publicado: (2026)
por: Li, Lei, et al.
Publicado: (2026)
DualMap: Online Open-Vocabulary Semantic Mapping for Natural Language Navigation in Dynamic Changing Scenes
por: Jiang, Jiajun, et al.
Publicado: (2025)
por: Jiang, Jiajun, et al.
Publicado: (2025)
Improving Dialogue Discourse Parsing through Discourse-aware Utterance Clarification
por: Fan, Yaxin, et al.
Publicado: (2025)
por: Fan, Yaxin, et al.
Publicado: (2025)
Efficient Multi-Instance Generation with Janus-Pro-Dirven Prompt Parsing
por: Qi, Fan, et al.
Publicado: (2025)
por: Qi, Fan, et al.
Publicado: (2025)
Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
por: Fang, Haipeng, et al.
Publicado: (2025)
por: Fang, Haipeng, et al.
Publicado: (2025)
Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing
por: Liu, Fuyuan, et al.
Publicado: (2026)
por: Liu, Fuyuan, et al.
Publicado: (2026)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
por: Song, Jiahe, et al.
Publicado: (2025)
por: Song, Jiahe, et al.
Publicado: (2025)
Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data
por: Guan, Zhong, et al.
Publicado: (2025)
por: Guan, Zhong, et al.
Publicado: (2025)
PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL Translation
por: Zhou, Wei, et al.
Publicado: (2025)
por: Zhou, Wei, et al.
Publicado: (2025)
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
por: Zhao, Zhiyuan, et al.
Publicado: (2024)
por: Zhao, Zhiyuan, et al.
Publicado: (2024)
Ejemplares similares
-
MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
por: Wang, Bin, et al.
Publicado: (2026) -
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
por: Niu, Junbo, et al.
Publicado: (2025) -
MinerU: An Open-Source Solution for Precise Document Content Extraction
por: Wang, Bin, et al.
Publicado: (2024) -
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
por: Dong, Hejun, et al.
Publicado: (2026) -
MoDora: Tree-Based Semi-Structured Document Analysis System
por: Xu, Bangrui, et al.
Publicado: (2026)