MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Limisiewicz, Tomasz, Blevins, Terra, Gonen, Hila, Ahia, Orevaoghene, Zettlemoyer, Luke
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929584624107520
author Limisiewicz, Tomasz
Blevins, Terra
Gonen, Hila
Ahia, Orevaoghene
Zettlemoyer, Luke
author_facet Limisiewicz, Tomasz
Blevins, Terra
Gonen, Hila
Ahia, Orevaoghene
Zettlemoyer, Luke
contents A major consideration in multilingual language modeling is how to best represent languages with diverse vocabularies and scripts. Although contemporary text encoding methods cover most of the world's writing systems, they exhibit bias towards the high-resource languages of the Global West. As a result, texts of underrepresented languages tend to be segmented into long sequences of linguistically meaningless units. To address the disparities, we introduce a new paradigm that encodes the same information with segments of consistent size across diverse languages. Our encoding convention (MYTE) is based on morphemes, as their inventories are more balanced across languages than characters, which are used in previous methods. We show that MYTE produces shorter encodings for all 99 analyzed languages, with the most notable improvements for non-European languages and non-Latin scripts. This, in turn, improves multilingual LM performance and diminishes the perplexity gap throughout diverse languages.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10691
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling
Limisiewicz, Tomasz
Blevins, Terra
Gonen, Hila
Ahia, Orevaoghene
Zettlemoyer, Luke
Computation and Language
Artificial Intelligence
Machine Learning
A major consideration in multilingual language modeling is how to best represent languages with diverse vocabularies and scripts. Although contemporary text encoding methods cover most of the world's writing systems, they exhibit bias towards the high-resource languages of the Global West. As a result, texts of underrepresented languages tend to be segmented into long sequences of linguistically meaningless units. To address the disparities, we introduce a new paradigm that encodes the same information with segments of consistent size across diverse languages. Our encoding convention (MYTE) is based on morphemes, as their inventories are more balanced across languages than characters, which are used in previous methods. We show that MYTE produces shorter encodings for all 99 analyzed languages, with the most notable improvements for non-European languages and non-Latin scripts. This, in turn, improves multilingual LM performance and diminishes the perplexity gap throughout diverse languages.
title MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2403.10691