Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Madhavi, Hrishit, Cherian, Jacob, Khamkar, Yuvraj, Bhagat, Dhananjay
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910947842457600
author Madhavi, Hrishit
Cherian, Jacob
Khamkar, Yuvraj
Bhagat, Dhananjay
author_facet Madhavi, Hrishit
Cherian, Jacob
Khamkar, Yuvraj
Bhagat, Dhananjay
contents This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical Character Recognition (Tesseract) to extract text in languages such as English, Hindi, and Tamil, and then a pipeline involving large language model APIs (Gemini) for cross-lingual translation, abstractive summarization, and re-translation into a target language. Additional modules add sentiment analysis (TensorFlow), topic classification (Transformers), and date extraction (Regex) for better document comprehension. Made available in an accessible Gradio interface, the current research shows a real-world application of libraries, models, and APIs to close the language gap and enhance access to information in image media across different linguistic environments
format Preprint
id arxiv_https___arxiv_org_abs_2505_11177
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline
Madhavi, Hrishit
Cherian, Jacob
Khamkar, Yuvraj
Bhagat, Dhananjay
Computation and Language
Artificial Intelligence
68T50 (Natural language processing), 68U10 (Image processing)
This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical Character Recognition (Tesseract) to extract text in languages such as English, Hindi, and Tamil, and then a pipeline involving large language model APIs (Gemini) for cross-lingual translation, abstractive summarization, and re-translation into a target language. Additional modules add sentiment analysis (TensorFlow), topic classification (Transformers), and date extraction (Regex) for better document comprehension. Made available in an accessible Gradio interface, the current research shows a real-world application of libraries, models, and APIs to close the language gap and enhance access to information in image media across different linguistic environments
title Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline
topic Computation and Language
Artificial Intelligence
68T50 (Natural language processing), 68U10 (Image processing)
url https://arxiv.org/abs/2505.11177