An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rusli, Andre, Shishido, Makoto
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913624920948736
author Rusli, Andre
Shishido, Makoto
author_facet Rusli, Andre
Shishido, Makoto
contents This paper analyses how traditional baseline metrics, such as BLEU and TER, and neural-based methods, such as BERTScore and COMET, score several NMT models performance on chat translation and how these metrics perform when compared to human-annotated scores. The results show that for ranking NMT models in chat translations, all metrics seem consistent in deciding which model outperforms the others. This implies that traditional baseline metrics, which are faster and simpler to use, can still be helpful. On the other hand, when it comes to better correlation with human judgment, neural-based metrics outperform traditional metrics, with COMET achieving the highest correlation with the human-annotated score on a chat translation. However, we show that even the best metric struggles when scoring English translations from sentences with anaphoric zero-pronoun in Japanese.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18190
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation
Rusli, Andre
Shishido, Makoto
Computation and Language
Artificial Intelligence
This paper analyses how traditional baseline metrics, such as BLEU and TER, and neural-based methods, such as BERTScore and COMET, score several NMT models performance on chat translation and how these metrics perform when compared to human-annotated scores. The results show that for ranking NMT models in chat translations, all metrics seem consistent in deciding which model outperforms the others. This implies that traditional baseline metrics, which are faster and simpler to use, can still be helpful. On the other hand, when it comes to better correlation with human judgment, neural-based metrics outperform traditional metrics, with COMET achieving the highest correlation with the human-annotated score on a chat translation. However, we show that even the best metric struggles when scoring English translations from sentences with anaphoric zero-pronoun in Japanese.
title An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.18190