DocParseNet: Advanced Semantic Segmentation and OCR Embeddings for Efficient Scanned Document Annotation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mohammadshirazi, Ahmad, Firoozsalari, Ali Nosrati, Zhou, Mengxi, Kulshrestha, Dheeraj, Ramnath, Rajiv |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025)
MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
InsightBuild: LLM-Powered Causal Reasoning in Smart Building Systems
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
PIAD-SRNN: Physics-Informed Adaptive Decomposition in State-Space RNN
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
ALCo-FM: Adaptive Long-Context Foundation Model for Accident Prediction
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
Multimodal OCR: Parse Anything from Documents
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
von: Zheng, Handong, et al.
Veröffentlicht: (2026)
Masked LoGoNet: Fast and Accurate 3D Image Analysis for Medical Domain
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
KnobGen: Controlling the Sophistication of Artwork in Sketch-Based Diffusion Models
von: Navard, Pouyan, et al.
Veröffentlicht: (2024)
von: Navard, Pouyan, et al.
Veröffentlicht: (2024)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
von: Tang, Zihan, et al.
Veröffentlicht: (2026)
von: Tang, Zihan, et al.
Veröffentlicht: (2026)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
von: Hu, Anwen, et al.
Veröffentlicht: (2024)
von: Hu, Anwen, et al.
Veröffentlicht: (2024)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
von: Li, Zhang, et al.
Veröffentlicht: (2025)
von: Li, Zhang, et al.
Veröffentlicht: (2025)
mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
von: Hu, Anwen, et al.
Veröffentlicht: (2024)
von: Hu, Anwen, et al.
Veröffentlicht: (2024)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
von: Zhou, Changda, et al.
Veröffentlicht: (2026)
von: Zhou, Changda, et al.
Veröffentlicht: (2026)
Controlla: Learning Controllability via Graph-Constrained Latent Geometry
von: Murthy, Jamuna S., et al.
Veröffentlicht: (2026)
von: Murthy, Jamuna S., et al.
Veröffentlicht: (2026)
TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
ESA: Annotation-Efficient Active Learning for Semantic Segmentation
von: Ge, Jinchao, et al.
Veröffentlicht: (2024)
von: Ge, Jinchao, et al.
Veröffentlicht: (2024)
Seeing Straight: Document Orientation Detection for Efficient OCR
von: Goswami, Suranjan, et al.
Veröffentlicht: (2025)
von: Goswami, Suranjan, et al.
Veröffentlicht: (2025)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
von: Al-Homoud, Haneen, et al.
Veröffentlicht: (2025)
von: Al-Homoud, Haneen, et al.
Veröffentlicht: (2025)
Efficient and Interpretable Information Retrieval for Product Question Answering with Heterogeneous Data
von: Biswas, Biplob, et al.
Veröffentlicht: (2024)
von: Biswas, Biplob, et al.
Veröffentlicht: (2024)
ProMi: An Efficient Prototype-Mixture Baseline for Few-Shot Segmentation with Bounding-Box Annotations
von: Chiaroni, Florent, et al.
Veröffentlicht: (2025)
von: Chiaroni, Florent, et al.
Veröffentlicht: (2025)
DocSynthv2: A Practical Autoregressive Modeling for Document Generation
von: Biswas, Sanket, et al.
Veröffentlicht: (2024)
von: Biswas, Sanket, et al.
Veröffentlicht: (2024)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
von: Wang, Baode, et al.
Veröffentlicht: (2025)
von: Wang, Baode, et al.
Veröffentlicht: (2025)
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval
von: Most, Alexander, et al.
Veröffentlicht: (2025)
von: Most, Alexander, et al.
Veröffentlicht: (2025)
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
von: Cui, Cheng, et al.
Veröffentlicht: (2026)
CrashFormer: A Multimodal Architecture to Predict the Risk of Crash
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
Enhancement of Bengali OCR by Specialized Models and Advanced Techniques for Diverse Document Types
von: Rabby, AKM Shahariar Azad, et al.
Veröffentlicht: (2024)
von: Rabby, AKM Shahariar Azad, et al.
Veröffentlicht: (2024)
OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
von: Zhang, Yulong
Veröffentlicht: (2025)
von: Zhang, Yulong
Veröffentlicht: (2025)
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
von: Cui, Cheng, et al.
Veröffentlicht: (2025)
Noisy Annotations in Semantic Segmentation
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
ERPA: Efficient RPA Model Integrating OCR and LLMs for Intelligent Document Processing
von: Abdellaif, Osama, et al.
Veröffentlicht: (2024)
von: Abdellaif, Osama, et al.
Veröffentlicht: (2024)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025) -
MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2025) -
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025) -
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024) -
InsightBuild: LLM-Powered Causal Reasoning in Smart Building Systems
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)