Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Batsuren, Khuyagbaatar, Vylomova, Ekaterina, Dankers, Verna, Delgerbaatar, Tsetsuukhei, Uzan, Omri, Pinter, Yuval, Bella, Gábor
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914762682531840
author Batsuren, Khuyagbaatar
Vylomova, Ekaterina
Dankers, Verna
Delgerbaatar, Tsetsuukhei
Uzan, Omri
Pinter, Yuval
Bella, Gábor
author_facet Batsuren, Khuyagbaatar
Vylomova, Ekaterina
Dankers, Verna
Delgerbaatar, Tsetsuukhei
Uzan, Omri
Pinter, Yuval
Bella, Gábor
contents The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance of the models. While many improved tokenization algorithms have been proposed, their evaluation and cross-comparison is still an open problem. As a solution, we propose a combined intrinsic-extrinsic evaluation framework for subword tokenization. Intrinsic evaluation is based on our new UniMorph Labeller tool that classifies subword tokenization as either morphological or alien. Extrinsic evaluation, in turn, is performed via the Out-of-Vocabulary Generalization Challenge 1.0 benchmark, which consists of three newly specified downstream text classification tasks. Our empirical findings show that the accuracy of UniMorph Labeller is 98%, and that, in all language models studied (including ALBERT, BERT, RoBERTa, and DeBERTa), alien tokenization leads to poorer generalizations compared to morphological tokenization for semantic compositionality of word meanings.
format Preprint
id arxiv_https___arxiv_org_abs_2404_13292
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
Batsuren, Khuyagbaatar
Vylomova, Ekaterina
Dankers, Verna
Delgerbaatar, Tsetsuukhei
Uzan, Omri
Pinter, Yuval
Bella, Gábor
Computation and Language
Artificial Intelligence
The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance of the models. While many improved tokenization algorithms have been proposed, their evaluation and cross-comparison is still an open problem. As a solution, we propose a combined intrinsic-extrinsic evaluation framework for subword tokenization. Intrinsic evaluation is based on our new UniMorph Labeller tool that classifies subword tokenization as either morphological or alien. Extrinsic evaluation, in turn, is performed via the Out-of-Vocabulary Generalization Challenge 1.0 benchmark, which consists of three newly specified downstream text classification tasks. Our empirical findings show that the accuracy of UniMorph Labeller is 98%, and that, in all language models studied (including ALBERT, BERT, RoBERTa, and DeBERTa), alien tokenization leads to poorer generalizations compared to morphological tokenization for semantic compositionality of word meanings.
title Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.13292