Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dang, Thao Anh, Raviv, Limor, Galke, Lukas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914982788071424
author Dang, Thao Anh
Raviv, Limor
Galke, Lukas
author_facet Dang, Thao Anh
Raviv, Limor
Galke, Lukas
contents Morphology is a crucial factor for multilingual language modeling as it poses direct challenges for tokenization. Here, we seek to understand how tokenization influences the morphological knowledge encoded in multilingual language models. Specifically, we capture the impact of tokenization by contrasting two multilingual language models: mT5 and ByT5. The two models share the same architecture, training objective, and training data and only differ in their tokenization strategies: subword tokenization vs.\@ character-level tokenization. Probing the morphological knowledge encoded in these models on four tasks and 17 languages, our analyses show that the models learn the morphological systems of some languages better than others and that morphological information is encoded in the middle and late layers. Finally, we show that languages with more irregularities benefit more from having a higher share of the pre-training data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11627
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
Dang, Thao Anh
Raviv, Limor
Galke, Lukas
Computation and Language
I.2.7
Morphology is a crucial factor for multilingual language modeling as it poses direct challenges for tokenization. Here, we seek to understand how tokenization influences the morphological knowledge encoded in multilingual language models. Specifically, we capture the impact of tokenization by contrasting two multilingual language models: mT5 and ByT5. The two models share the same architecture, training objective, and training data and only differ in their tokenization strategies: subword tokenization vs.\@ character-level tokenization. Probing the morphological knowledge encoded in these models on four tasks and 17 languages, our analyses show that the models learn the morphological systems of some languages better than others and that morphological information is encoded in the middle and late layers. Finally, we show that languages with more irregularities benefit more from having a higher share of the pre-training data.
title Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2410.11627