Towards Deployable OCR models for Indic languages
Fuente:
arXiv
Guardado en:
| Autores principales: | Mathew, Minesh, Mondal, Ajoy, Jawahar, CV |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
por: Pal, Aniket, et al.
Publicado: (2025)
por: Pal, Aniket, et al.
Publicado: (2025)
IndicSTR12: A Dataset for Indic Scene Text Recognition
por: Lunia, Harsh, et al.
Publicado: (2024)
por: Lunia, Harsh, et al.
Publicado: (2024)
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
por: Pal, Aniket, et al.
Publicado: (2024)
por: Pal, Aniket, et al.
Publicado: (2024)
Reading Between the Lanes: Text VideoQA on the Road
por: Tom, George, et al.
Publicado: (2023)
por: Tom, George, et al.
Publicado: (2023)
olmOCR 2: Unit Test Rewards for Document OCR
por: Poznanski, Jake, et al.
Publicado: (2025)
por: Poznanski, Jake, et al.
Publicado: (2025)
Navigating Text-to-Image Generative Bias across Indic Languages
por: Mittal, Surbhi, et al.
Publicado: (2024)
por: Mittal, Surbhi, et al.
Publicado: (2024)
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
por: Maniyar, Suyash, et al.
Publicado: (2025)
por: Maniyar, Suyash, et al.
Publicado: (2025)
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy
por: Guan, Shuhao, et al.
Publicado: (2025)
por: Guan, Shuhao, et al.
Publicado: (2025)
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts
por: Kargaran, Amir Hossein, et al.
Publicado: (2026)
por: Kargaran, Amir Hossein, et al.
Publicado: (2026)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
por: Gagnier, Henry, et al.
Publicado: (2026)
por: Gagnier, Henry, et al.
Publicado: (2026)
Improving OCR for Historical Texts of Multiple Languages
por: Westerdijk, Hylke, et al.
Publicado: (2025)
por: Westerdijk, Hylke, et al.
Publicado: (2025)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
por: Kashid, Harshvivek, et al.
Publicado: (2024)
por: Kashid, Harshvivek, et al.
Publicado: (2024)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
por: Liu, Yuliang, et al.
Publicado: (2023)
por: Liu, Yuliang, et al.
Publicado: (2023)
Seeing Straight: Document Orientation Detection for Efficient OCR
por: Goswami, Suranjan, et al.
Publicado: (2025)
por: Goswami, Suranjan, et al.
Publicado: (2025)
ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding
por: Abdallah, Abdelrahman, et al.
Publicado: (2024)
por: Abdallah, Abdelrahman, et al.
Publicado: (2024)
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
por: Hennara, Khalil, et al.
Publicado: (2025)
por: Hennara, Khalil, et al.
Publicado: (2025)
Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
por: Liang, Yunhao, et al.
Publicado: (2026)
por: Liang, Yunhao, et al.
Publicado: (2026)
Advances and Limitations in Open Source Arabic-Script OCR: A Case Study
por: Kiessling, Benjamin, et al.
Publicado: (2024)
por: Kiessling, Benjamin, et al.
Publicado: (2024)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
por: Wang, Zhengren, et al.
Publicado: (2026)
por: Wang, Zhengren, et al.
Publicado: (2026)
Vision Meets Language: A RAG-Augmented YOLOv8 Framework for Coffee Disease Diagnosis and Farmer Assistance
por: Mondal, Semanto
Publicado: (2025)
por: Mondal, Semanto
Publicado: (2025)
DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
por: Cardoso, Gabriel Pimenta de Freitas, et al.
Publicado: (2026)
por: Cardoso, Gabriel Pimenta de Freitas, et al.
Publicado: (2026)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
por: Do, Thao, et al.
Publicado: (2024)
por: Do, Thao, et al.
Publicado: (2024)
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
por: Tang, Zihan, et al.
Publicado: (2026)
por: Tang, Zihan, et al.
Publicado: (2026)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
por: Ye, Maoyuan, et al.
Publicado: (2025)
por: Ye, Maoyuan, et al.
Publicado: (2025)
DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenarios
por: Artham, Sainithin, et al.
Publicado: (2026)
por: Artham, Sainithin, et al.
Publicado: (2026)
Frame Sampling Strategies Matter: A Benchmark for small vision language models
por: Brkic, Marija, et al.
Publicado: (2025)
por: Brkic, Marija, et al.
Publicado: (2025)
Confidence-Aware Document OCR Error Detection
por: Hemmer, Arthur, et al.
Publicado: (2024)
por: Hemmer, Arthur, et al.
Publicado: (2024)
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding
por: Le, Binh M., et al.
Publicado: (2025)
por: Le, Binh M., et al.
Publicado: (2025)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
por: Al-Homoud, Haneen, et al.
Publicado: (2025)
por: Al-Homoud, Haneen, et al.
Publicado: (2025)
PubMed-OCR: PMC Open Access OCR Annotations
por: Heidenreich, Hunter, et al.
Publicado: (2026)
por: Heidenreich, Hunter, et al.
Publicado: (2026)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
por: Tan, Alvin Wei Ming, et al.
Publicado: (2025)
por: Tan, Alvin Wei Ming, et al.
Publicado: (2025)
ProfVLM: A lightweight video-language model for multi-view proficiency estimation
por: Bianchi, Edoardo, et al.
Publicado: (2025)
por: Bianchi, Edoardo, et al.
Publicado: (2025)
PubMed-Ophtha: An open resource for training ophthalmology vision-language models on scientific literature
por: Hallitschke, Verena Jasmin, et al.
Publicado: (2026)
por: Hallitschke, Verena Jasmin, et al.
Publicado: (2026)
Interpreting the linear structure of vision-language model embedding spaces
por: Papadimitriou, Isabel, et al.
Publicado: (2025)
por: Papadimitriou, Isabel, et al.
Publicado: (2025)
Image captioning in different languages
por: van Miltenburg, Emiel
Publicado: (2024)
por: van Miltenburg, Emiel
Publicado: (2024)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
por: Zhong, Yufeng, et al.
Publicado: (2025)
por: Zhong, Yufeng, et al.
Publicado: (2025)
SmoGVLM: A Small, Graph-enhanced Vision-Language Model
por: Mondal, Debjyoti, et al.
Publicado: (2026)
por: Mondal, Debjyoti, et al.
Publicado: (2026)
Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction
por: Rashad, Mohamed
Publicado: (2024)
por: Rashad, Mohamed
Publicado: (2024)
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
por: Salazar, Israfel, et al.
Publicado: (2025)
por: Salazar, Israfel, et al.
Publicado: (2025)
Superhuman performance in urology board questions by an explainable large language model enabled for context integration of the European Association of Urology guidelines: the UroBot study
por: Hetz, Martin J., et al.
Publicado: (2024)
por: Hetz, Martin J., et al.
Publicado: (2024)
Ejemplares similares
-
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
por: Pal, Aniket, et al.
Publicado: (2025) -
IndicSTR12: A Dataset for Indic Scene Text Recognition
por: Lunia, Harsh, et al.
Publicado: (2024) -
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
por: Pal, Aniket, et al.
Publicado: (2024) -
Reading Between the Lanes: Text VideoQA on the Road
por: Tom, George, et al.
Publicado: (2023) -
olmOCR 2: Unit Test Rewards for Document OCR
por: Poznanski, Jake, et al.
Publicado: (2025)