Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rezaeimanesh, Sara, Hosseini, Faezeh, Yaghoobzadeh, Yadollah
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908237854408704
author Rezaeimanesh, Sara
Hosseini, Faezeh
Yaghoobzadeh, Yadollah
author_facet Rezaeimanesh, Sara
Hosseini, Faezeh
Yaghoobzadeh, Yadollah
contents Large language models (LLMs) have shown superior capabilities in translating figurative language compared to neural machine translation (NMT) systems. However, the impact of different prompting methods and LLM-NMT combinations on idiom translation has yet to be thoroughly investigated. This paper introduces two parallel datasets of sentences containing idiomatic expressions for Persian$\rightarrow$English and English$\rightarrow$Persian translations, with Persian idioms sampled from our PersianIdioms resource, a collection of 2,200 idioms and their meanings, with 700 including usage examples. Using these datasets, we evaluate various open- and closed-source LLMs, NMT models, and their combinations. Translation quality is assessed through idiom translation accuracy and fluency. We also find that automatic evaluation methods like LLM-as-a-judge, BLEU, and BERTScore are effective for comparing different aspects of model performance. Our experiments reveal that Claude-3.5-Sonnet delivers outstanding results in both translation directions. For English$\rightarrow$Persian, combining weaker LLMs with Google Translate improves results, while Persian$\rightarrow$English translations benefit from single prompts for simpler models and complex prompts for advanced ones.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09993
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation
Rezaeimanesh, Sara
Hosseini, Faezeh
Yaghoobzadeh, Yadollah
Computation and Language
I.2.7
Large language models (LLMs) have shown superior capabilities in translating figurative language compared to neural machine translation (NMT) systems. However, the impact of different prompting methods and LLM-NMT combinations on idiom translation has yet to be thoroughly investigated. This paper introduces two parallel datasets of sentences containing idiomatic expressions for Persian$\rightarrow$English and English$\rightarrow$Persian translations, with Persian idioms sampled from our PersianIdioms resource, a collection of 2,200 idioms and their meanings, with 700 including usage examples. Using these datasets, we evaluate various open- and closed-source LLMs, NMT models, and their combinations. Translation quality is assessed through idiom translation accuracy and fluency. We also find that automatic evaluation methods like LLM-as-a-judge, BLEU, and BERTScore are effective for comparing different aspects of model performance. Our experiments reveal that Claude-3.5-Sonnet delivers outstanding results in both translation directions. For English$\rightarrow$Persian, combining weaker LLMs with Google Translate improves results, while Persian$\rightarrow$English translations benefit from single prompts for simpler models and complex prompts for advanced ones.
title Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2412.09993