Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zaranis, Emmanouil, Attanasio, Giuseppe, Agrawal, Sweta, Martins, André F. T.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912409922306048
author Zaranis, Emmanouil
Attanasio, Giuseppe
Agrawal, Sweta
Martins, André F. T.
author_facet Zaranis, Emmanouil
Attanasio, Giuseppe
Agrawal, Sweta
Martins, André F. T.
contents Quality estimation (QE)-the automatic assessment of translation quality-has recently become crucial across several stages of the translation pipeline, from data curation to training and decoding. While QE metrics have been optimized to align with human judgments, whether they encode social biases has been largely overlooked. Biased QE risks favoring certain demographic groups over others, e.g., by exacerbating gaps in visibility and usability. This paper defines and investigates gender bias of QE metrics and discusses its downstream implications for machine translation (MT). Experiments with state-of-the-art QE metrics across multiple domains, datasets, and languages reveal significant bias. When a human entity's gender in the source is undisclosed, masculine-inflected translations score higher than feminine-inflected ones, and gender-neutral translations are penalized. Even when contextual cues disambiguate gender, using context-aware QE metrics leads to more errors in selecting the correct translation inflection for feminine referents than for masculine ones. Moreover, a biased QE metric affects data filtering and quality-aware decoding. Our findings underscore the need for a renewed focus on developing and evaluating QE metrics centered on gender.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10995
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
Zaranis, Emmanouil
Attanasio, Giuseppe
Agrawal, Sweta
Martins, André F. T.
Computation and Language
Quality estimation (QE)-the automatic assessment of translation quality-has recently become crucial across several stages of the translation pipeline, from data curation to training and decoding. While QE metrics have been optimized to align with human judgments, whether they encode social biases has been largely overlooked. Biased QE risks favoring certain demographic groups over others, e.g., by exacerbating gaps in visibility and usability. This paper defines and investigates gender bias of QE metrics and discusses its downstream implications for machine translation (MT). Experiments with state-of-the-art QE metrics across multiple domains, datasets, and languages reveal significant bias. When a human entity's gender in the source is undisclosed, masculine-inflected translations score higher than feminine-inflected ones, and gender-neutral translations are penalized. Even when contextual cues disambiguate gender, using context-aware QE metrics leads to more errors in selecting the correct translation inflection for feminine referents than for masculine ones. Moreover, a biased QE metric affects data filtering and quality-aware decoding. Our findings underscore the need for a renewed focus on developing and evaluating QE metrics centered on gender.
title Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
topic Computation and Language
url https://arxiv.org/abs/2410.10995