Saved in:
Bibliographic Details
Main Author: Yanampally, Abhiram Reddy
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.05914
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916678923714560
author Yanampally, Abhiram Reddy
author_facet Yanampally, Abhiram Reddy
contents This paper presents a novel approach to constructing an English-to-Telugu translation model by leveraging transfer learning techniques and addressing the challenges associated with low-resource languages. Utilizing the Bharat Parallel Corpus Collection (BPCC) as the primary dataset, the model incorporates iterative backtranslation to generate synthetic parallel data, effectively augmenting the training dataset and enhancing the model's translation capabilities. The research focuses on a comprehensive strategy for improving model performance through data augmentation, optimization of training parameters, and the effective use of pre-trained models. These methodologies aim to create a robust translation system that can handle diverse sentence structures and linguistic nuances in both English and Telugu. This work highlights the significance of innovative data handling techniques and the potential of transfer learning in overcoming limitations posed by sparse datasets in low-resource languages. The study contributes to the field of machine translation and seeks to improve communication between English and Telugu speakers in practical contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05914
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle High-Resource Translation:Turning Abundance into Accessibility
Yanampally, Abhiram Reddy
Computation and Language
This paper presents a novel approach to constructing an English-to-Telugu translation model by leveraging transfer learning techniques and addressing the challenges associated with low-resource languages. Utilizing the Bharat Parallel Corpus Collection (BPCC) as the primary dataset, the model incorporates iterative backtranslation to generate synthetic parallel data, effectively augmenting the training dataset and enhancing the model's translation capabilities. The research focuses on a comprehensive strategy for improving model performance through data augmentation, optimization of training parameters, and the effective use of pre-trained models. These methodologies aim to create a robust translation system that can handle diverse sentence structures and linguistic nuances in both English and Telugu. This work highlights the significance of innovative data handling techniques and the potential of transfer learning in overcoming limitations posed by sparse datasets in low-resource languages. The study contributes to the field of machine translation and seeks to improve communication between English and Telugu speakers in practical contexts.
title High-Resource Translation:Turning Abundance into Accessibility
topic Computation and Language
url https://arxiv.org/abs/2504.05914