Distributional Properties of Subword Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cognetta, Marco, Zouhar, Vilém, Okazaki, Naoaki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911997698768896
author Cognetta, Marco
Zouhar, Vilém
Okazaki, Naoaki
author_facet Cognetta, Marco
Zouhar, Vilém
Okazaki, Naoaki
contents Subword regularization, used widely in NLP, improves model performance by reducing the dependency on exact tokenizations, augmenting the training corpus, and exposing the model to more unique contexts during training. BPE and MaxMatch, two popular subword tokenization schemes, have stochastic dropout regularization variants. However, there has not been an analysis of the distributions formed by them. We show that these stochastic variants are heavily biased towards a small set of tokenizations per word. If the benefits of subword regularization are as mentioned, we hypothesize that biasedness artificially limits the effectiveness of these schemes. Thus, we propose an algorithm to uniformly sample tokenizations that we use as a drop-in replacement for the stochastic aspects of existing tokenizers, and find that it improves machine translation quality.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11443
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Distributional Properties of Subword Regularization
Cognetta, Marco
Zouhar, Vilém
Okazaki, Naoaki
Computation and Language
Subword regularization, used widely in NLP, improves model performance by reducing the dependency on exact tokenizations, augmenting the training corpus, and exposing the model to more unique contexts during training. BPE and MaxMatch, two popular subword tokenization schemes, have stochastic dropout regularization variants. However, there has not been an analysis of the distributions formed by them. We show that these stochastic variants are heavily biased towards a small set of tokenizations per word. If the benefits of subword regularization are as mentioned, we hypothesize that biasedness artificially limits the effectiveness of these schemes. Thus, we propose an algorithm to uniformly sample tokenizations that we use as a drop-in replacement for the stochastic aspects of existing tokenizers, and find that it improves machine translation quality.
title Distributional Properties of Subword Regularization
topic Computation and Language
url https://arxiv.org/abs/2408.11443