Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910947842457600 |
|---|---|
| author | Madhavi, Hrishit Cherian, Jacob Khamkar, Yuvraj Bhagat, Dhananjay |
| author_facet | Madhavi, Hrishit Cherian, Jacob Khamkar, Yuvraj Bhagat, Dhananjay |
| contents | This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical Character Recognition (Tesseract) to extract text in languages such as English, Hindi, and Tamil, and then a pipeline involving large language model APIs (Gemini) for cross-lingual translation, abstractive summarization, and re-translation into a target language. Additional modules add sentiment analysis (TensorFlow), topic classification (Transformers), and date extraction (Regex) for better document comprehension. Made available in an accessible Gradio interface, the current research shows a real-world application of libraries, models, and APIs to close the language gap and enhance access to information in image media across different linguistic environments |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_11177 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline Madhavi, Hrishit Cherian, Jacob Khamkar, Yuvraj Bhagat, Dhananjay Computation and Language Artificial Intelligence 68T50 (Natural language processing), 68U10 (Image processing) This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical Character Recognition (Tesseract) to extract text in languages such as English, Hindi, and Tamil, and then a pipeline involving large language model APIs (Gemini) for cross-lingual translation, abstractive summarization, and re-translation into a target language. Additional modules add sentiment analysis (TensorFlow), topic classification (Transformers), and date extraction (Regex) for better document comprehension. Made available in an accessible Gradio interface, the current research shows a real-world application of libraries, models, and APIs to close the language gap and enhance access to information in image media across different linguistic environments |
| title | Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline |
| topic | Computation and Language Artificial Intelligence 68T50 (Natural language processing), 68U10 (Image processing) |
| url | https://arxiv.org/abs/2505.11177 |