On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Briakou, Eleftheria, Liu, Zhongtao, Cherry, Colin, Freitag, Markus
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916419071901696
author Briakou, Eleftheria
Liu, Zhongtao
Cherry, Colin
Freitag, Markus
author_facet Briakou, Eleftheria
Liu, Zhongtao
Cherry, Colin
Freitag, Markus
contents This paper investigates the impact of verbose LLM translations on evaluation. We first demonstrate the prevalence of this behavior across several LLM outputs drawn from the WMT 2024 general shared task on machine translation. We then identify the primary triggers of verbosity, including safety, copyright concerns, and insufficient context in short input queries. Finally, we show that ignoring this behavior unfairly penalizes more verbose LLMs according to both automatic and human evaluations, highlighting the need to address this issue for more accurate future evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00863
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
Briakou, Eleftheria
Liu, Zhongtao
Cherry, Colin
Freitag, Markus
Computation and Language
This paper investigates the impact of verbose LLM translations on evaluation. We first demonstrate the prevalence of this behavior across several LLM outputs drawn from the WMT 2024 general shared task on machine translation. We then identify the primary triggers of verbosity, including safety, copyright concerns, and insufficient context in short input queries. Finally, we show that ignoring this behavior unfairly penalizes more verbose LLMs according to both automatic and human evaluations, highlighting the need to address this issue for more accurate future evaluations.
title On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
topic Computation and Language
url https://arxiv.org/abs/2410.00863