Misgendering and Assuming Gender in Machine Translation when Working with Low-Resource Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghosh, Sourojit, Chatterjee, Srishti
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917611133992960
author Ghosh, Sourojit
Chatterjee, Srishti
author_facet Ghosh, Sourojit
Chatterjee, Srishti
contents This chapter focuses on gender-related errors in machine translation (MT) in the context of low-resource languages. We begin by explaining what low-resource languages are, examining the inseparable social and computational factors that create such linguistic hierarchies. We demonstrate through a case study of our mother tongue Bengali, a global language spoken by almost 300 million people but still classified as low-resource, how gender is assumed and inferred in translations to and from the high(est)-resource English when no such information is provided in source texts. We discuss the postcolonial and societal impacts of such errors leading to linguistic erasure and representational harms, and conclude by discussing potential solutions towards uplifting languages by providing them more agency in MT conversations.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13165
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Misgendering and Assuming Gender in Machine Translation when Working with Low-Resource Languages
Ghosh, Sourojit
Chatterjee, Srishti
Computation and Language
This chapter focuses on gender-related errors in machine translation (MT) in the context of low-resource languages. We begin by explaining what low-resource languages are, examining the inseparable social and computational factors that create such linguistic hierarchies. We demonstrate through a case study of our mother tongue Bengali, a global language spoken by almost 300 million people but still classified as low-resource, how gender is assumed and inferred in translations to and from the high(est)-resource English when no such information is provided in source texts. We discuss the postcolonial and societal impacts of such errors leading to linguistic erasure and representational harms, and conclude by discussing potential solutions towards uplifting languages by providing them more agency in MT conversations.
title Misgendering and Assuming Gender in Machine Translation when Working with Low-Resource Languages
topic Computation and Language
url https://arxiv.org/abs/2401.13165