Optimizing Nepali PDF Extraction: A Comparative Study of Parser and OCR Technologies
Fuente:
arXiv
Guardado en:
| Autores principales: | Paudel, Prabin, Khadka, Supriya, C., Ranju G., Shah, Rahul, Joshi, Basanta |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models
por: Jin, Yichao, et al.
Publicado: (2025)
por: Jin, Yichao, et al.
Publicado: (2025)
Evidence Units: Ontology-Grounded Document Organization for Parser-Independent Retrieval
por: Han, Yeonjee
Publicado: (2026)
por: Han, Yeonjee
Publicado: (2026)
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
por: Horn, Pius, et al.
Publicado: (2025)
por: Horn, Pius, et al.
Publicado: (2025)
Digitization of Document and Information Extraction using OCR
por: Sinha, Rasha, et al.
Publicado: (2025)
por: Sinha, Rasha, et al.
Publicado: (2025)
Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications
por: Begha, Funghang Limbu, et al.
Publicado: (2026)
por: Begha, Funghang Limbu, et al.
Publicado: (2026)
Content-based Recommendation Engine for Video Streaming Platform
por: Khadka, Puskal, et al.
Publicado: (2023)
por: Khadka, Puskal, et al.
Publicado: (2023)
A Comparative Study of PDF Parsing Tools Across Diverse Document Categories
por: Adhikari, Narayan S., et al.
Publicado: (2024)
por: Adhikari, Narayan S., et al.
Publicado: (2024)
Compact Multimodal Language Models as Robust OCR Alternatives for Noisy Textual Clinical Reports
por: Neveditsin, Nikita, et al.
Publicado: (2025)
por: Neveditsin, Nikita, et al.
Publicado: (2025)
Correlating Power Outage Spread with Infrastructure Interdependencies During Hurricanes
por: Bose, Avishek, et al.
Publicado: (2024)
por: Bose, Avishek, et al.
Publicado: (2024)
Narrative Trails: A Method for Coherent Storyline Extraction via Maximum Capacity Path Optimization
por: German, Fausto, et al.
Publicado: (2025)
por: German, Fausto, et al.
Publicado: (2025)
KAP: MLLM-assisted OCR Text Enhancement for Hybrid Retrieval in Chinese Non-Narrative Documents
por: Hsu, Hsin-Ling, et al.
Publicado: (2025)
por: Hsu, Hsin-Ling, et al.
Publicado: (2025)
Tabular PDF Information Extraction with Local LLMs and Layout-Aware Parsing: A Reliability Evaluation
por: Hilmi, Muhammad Anis Al, et al.
Publicado: (2026)
por: Hilmi, Muhammad Anis Al, et al.
Publicado: (2026)
Recommendation Algorithms: A Comparative Study in Movie Domain
por: Chivukula, Rohit, et al.
Publicado: (2026)
por: Chivukula, Rohit, et al.
Publicado: (2026)
Efficiency Optimizations for Superblock-based Sparse Retrieval
por: Carlson, Parker, et al.
Publicado: (2026)
por: Carlson, Parker, et al.
Publicado: (2026)
Beyond String Matching: Semantic Evaluation of PDF Table Extraction
por: Horn, Pius, et al.
Publicado: (2026)
por: Horn, Pius, et al.
Publicado: (2026)
Author Unknown: Evaluating Performance of Author Extraction Libraries on Global Online News Articles
por: Hatwar, Sriharsha, et al.
Publicado: (2024)
por: Hatwar, Sriharsha, et al.
Publicado: (2024)
A Comparative Study of Retrieval Methods in Azure AI Search
por: Mao, Qiang, et al.
Publicado: (2025)
por: Mao, Qiang, et al.
Publicado: (2025)
Optimizing RAG Techniques for Automotive Industry PDF Chatbots: A Case Study with Locally Deployed Ollama Models
por: Liu, Fei, et al.
Publicado: (2024)
por: Liu, Fei, et al.
Publicado: (2024)
A Systematic Replicability and Comparative Study of BSARec and SASRec for Sequential Recommendation
por: D'Ercoli, Chiara, et al.
Publicado: (2025)
por: D'Ercoli, Chiara, et al.
Publicado: (2025)
DocGraphLM: Documental Graph Language Model for Information Extraction
por: Wang, Dongsheng, et al.
Publicado: (2024)
por: Wang, Dongsheng, et al.
Publicado: (2024)
Information Extraction From Fiscal Documents Using LLMs
por: Aggarwal, Vikram, et al.
Publicado: (2025)
por: Aggarwal, Vikram, et al.
Publicado: (2025)
Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion
por: Sun, Huatuan, et al.
Publicado: (2025)
por: Sun, Huatuan, et al.
Publicado: (2025)
FABULA: Intelligence Report Generation Using Retrieval-Augmented Narrative Construction
por: Ranade, Priyanka, et al.
Publicado: (2023)
por: Ranade, Priyanka, et al.
Publicado: (2023)
Comparison of Feature Learning Methods for Metadata Extraction from PDF Scholarly Documents
por: Boukhers, Zeyd, et al.
Publicado: (2025)
por: Boukhers, Zeyd, et al.
Publicado: (2025)
Fundus: A Simple-to-Use News Scraper Optimized for High Quality Extractions
por: Dallabetta, Max, et al.
Publicado: (2024)
por: Dallabetta, Max, et al.
Publicado: (2024)
Retracted Citations and Self-citations in Retracted Publications: A Comparative Study of Plagiarism and Fake Peer Review
por: Sharmaa, Kiran, et al.
Publicado: (2025)
por: Sharmaa, Kiran, et al.
Publicado: (2025)
New Method for Keyword Extraction for Patent Claims
por: Rossi, Julien
Publicado: (2024)
por: Rossi, Julien
Publicado: (2024)
Comparative Analysis of Lion and AdamW Optimizers for Cross-Encoder Reranking with MiniLM, GTE, and ModernBERT
por: Kumar, Shahil, et al.
Publicado: (2025)
por: Kumar, Shahil, et al.
Publicado: (2025)
Scaling Automatic Extraction of Pseudocode
por: Toksoz, Levent, et al.
Publicado: (2024)
por: Toksoz, Levent, et al.
Publicado: (2024)
Uncertainty-Aware Complex Scientific Table Data Extraction
por: Ajayi, Kehinde, et al.
Publicado: (2025)
por: Ajayi, Kehinde, et al.
Publicado: (2025)
TRUE: A Reproducible Framework for LLM-Driven Relevance Judgment in Information Retrieval
por: Dewan, Mouly, et al.
Publicado: (2025)
por: Dewan, Mouly, et al.
Publicado: (2025)
Evaluating LLM Abilities to Understand Tabular Electronic Health Records: A Comprehensive Study of Patient Data Extraction and Retrieval
por: Lovon, Jesus, et al.
Publicado: (2025)
por: Lovon, Jesus, et al.
Publicado: (2025)
Information Extraction from Historical Well Records Using A Large Language Model
por: Ma, Zhiwei, et al.
Publicado: (2024)
por: Ma, Zhiwei, et al.
Publicado: (2024)
Decoding MIE: A Novel Dataset Approach Using Topic Extraction and Affiliation Parsing
por: Bitaraf, Ehsan, et al.
Publicado: (2024)
por: Bitaraf, Ehsan, et al.
Publicado: (2024)
Multi-Label Zero-Shot Product Attribute-Value Extraction
por: Gong, Jiaying, et al.
Publicado: (2024)
por: Gong, Jiaying, et al.
Publicado: (2024)
DELM: a Python toolkit for Data Extraction with Language Models
por: Fithian, Eric, et al.
Publicado: (2025)
por: Fithian, Eric, et al.
Publicado: (2025)
Reproducible Synthetic Clinical Letters for Seizure Frequency Information Extraction
por: Gan, Yujian, et al.
Publicado: (2026)
por: Gan, Yujian, et al.
Publicado: (2026)
RELATE: Relation Extraction in Biomedical Abstracts with LLMs and Ontology Constraints
por: Olasunkanmi, Olawumi, et al.
Publicado: (2025)
por: Olasunkanmi, Olawumi, et al.
Publicado: (2025)
Enhanced NIRMAL Optimizer With Damped Nesterov Acceleration: A Comparative Analysis
por: Gaud, Nirmal, et al.
Publicado: (2025)
por: Gaud, Nirmal, et al.
Publicado: (2025)
Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG
por: Bachyr, Omar El, et al.
Publicado: (2026)
por: Bachyr, Omar El, et al.
Publicado: (2026)
Ejemplares similares
-
Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models
por: Jin, Yichao, et al.
Publicado: (2025) -
Evidence Units: Ontology-Grounded Document Organization for Parser-Independent Retrieval
por: Han, Yeonjee
Publicado: (2026) -
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
por: Horn, Pius, et al.
Publicado: (2025) -
Digitization of Document and Information Extraction using OCR
por: Sinha, Rasha, et al.
Publicado: (2025) -
Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications
por: Begha, Funghang Limbu, et al.
Publicado: (2026)