Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Briva-Iglesias, Vicent, Camargo, Joao Lucas Cavalheiro, Dogru, Gokhan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2402.07681
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914675328811008
author Briva-Iglesias, Vicent
Camargo, Joao Lucas Cavalheiro
Dogru, Gokhan
author_facet Briva-Iglesias, Vicent
Camargo, Joao Lucas Cavalheiro
Dogru, Gokhan
contents This study evaluates the machine translation (MT) quality of two state-of-the-art large language models (LLMs) against a tradition-al neural machine translation (NMT) system across four language pairs in the legal domain. It combines automatic evaluation met-rics (AEMs) and human evaluation (HE) by professional transla-tors to assess translation ranking, fluency and adequacy. The re-sults indicate that while Google Translate generally outperforms LLMs in AEMs, human evaluators rate LLMs, especially GPT-4, comparably or slightly better in terms of producing contextually adequate and fluent translations. This discrepancy suggests LLMs' potential in handling specialized legal terminology and context, highlighting the importance of human evaluation methods in assessing MT quality. The study underscores the evolving capabil-ities of LLMs in specialized domains and calls for reevaluation of traditional AEMs to better capture the nuances of LLM-generated translations.
format Preprint
id arxiv_https___arxiv_org_abs_2402_07681
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Models "Ad Referendum": How Good Are They at Machine Translation in the Legal Domain?
Briva-Iglesias, Vicent
Camargo, Joao Lucas Cavalheiro
Dogru, Gokhan
Computation and Language
Artificial Intelligence
This study evaluates the machine translation (MT) quality of two state-of-the-art large language models (LLMs) against a tradition-al neural machine translation (NMT) system across four language pairs in the legal domain. It combines automatic evaluation met-rics (AEMs) and human evaluation (HE) by professional transla-tors to assess translation ranking, fluency and adequacy. The re-sults indicate that while Google Translate generally outperforms LLMs in AEMs, human evaluators rate LLMs, especially GPT-4, comparably or slightly better in terms of producing contextually adequate and fluent translations. This discrepancy suggests LLMs' potential in handling specialized legal terminology and context, highlighting the importance of human evaluation methods in assessing MT quality. The study underscores the evolving capabil-ities of LLMs in specialized domains and calls for reevaluation of traditional AEMs to better capture the nuances of LLM-generated translations.
title Large Language Models "Ad Referendum": How Good Are They at Machine Translation in the Legal Domain?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.07681