Reconsidering SMT Over NMT for Closely Related Languages: A Case Study of Persian-Hindi Pair

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yousofi, Waisullah, Bhattacharyya, Pushpak
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915075282960384
author Yousofi, Waisullah
Bhattacharyya, Pushpak
author_facet Yousofi, Waisullah
Bhattacharyya, Pushpak
contents This paper demonstrates that Phrase-Based Statistical Machine Translation (PBSMT) can outperform Transformer-based Neural Machine Translation (NMT) in moderate-resource scenarios, specifically for structurally similar languages, like the Persian-Hindi pair. Despite the Transformer architecture's typical preference for large parallel corpora, our results show that PBSMT achieves a BLEU score of 66.32, significantly exceeding the Transformer-NMT score of 53.7 on the same dataset. Additionally, we explore variations of the SMT architecture, including training on Romanized text and modifying the word order of Persian sentences to match the left-to-right (LTR) structure of Hindi. Our findings highlight the importance of choosing the right architecture based on language pair characteristics and advocate for SMT as a high-performing alternative, even in contexts commonly dominated by NMT.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16877
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reconsidering SMT Over NMT for Closely Related Languages: A Case Study of Persian-Hindi Pair
Yousofi, Waisullah
Bhattacharyya, Pushpak
Computation and Language
This paper demonstrates that Phrase-Based Statistical Machine Translation (PBSMT) can outperform Transformer-based Neural Machine Translation (NMT) in moderate-resource scenarios, specifically for structurally similar languages, like the Persian-Hindi pair. Despite the Transformer architecture's typical preference for large parallel corpora, our results show that PBSMT achieves a BLEU score of 66.32, significantly exceeding the Transformer-NMT score of 53.7 on the same dataset. Additionally, we explore variations of the SMT architecture, including training on Romanized text and modifying the word order of Persian sentences to match the left-to-right (LTR) structure of Hindi. Our findings highlight the importance of choosing the right architecture based on language pair characteristics and advocate for SMT as a high-performing alternative, even in contexts commonly dominated by NMT.
title Reconsidering SMT Over NMT for Closely Related Languages: A Case Study of Persian-Hindi Pair
topic Computation and Language
url https://arxiv.org/abs/2412.16877