Saved in:
Bibliographic Details
Main Authors: Lin, Pin-Jie, Chang, Ernie, Shi, Yangyang, Chandra, Vikas
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.13837
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912302123450368
author Lin, Pin-Jie
Chang, Ernie
Shi, Yangyang
Chandra, Vikas
author_facet Lin, Pin-Jie
Chang, Ernie
Shi, Yangyang
Chandra, Vikas
contents Past vocabulary learning techniques identify relevant vocabulary before training, relying on statistical and entropy-based assumptions that largely neglect the role of model training. Empirically, we observe that trained translation models are induced to use a byte-pair encoding (BPE) vocabulary subset distinct from the original BPE vocabulary, leading to performance improvements when retrained with the induced vocabulary. In this paper, we analyze this discrepancy in neural machine translation by examining vocabulary and entropy shifts during self-training--where each iteration generates a labeled dataset by pairing source sentences with the model's predictions to define a new vocabulary. Building on these insights, we propose self-vocabularizing training, an iterative method that self-selects a smaller, more optimal vocabulary, yielding up to a 1.49 BLEU improvement. Moreover, we find that deeper model architectures lead to both an increase in unique token usage and a 6-8% reduction in vocabulary size.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13837
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Vocabularizing Training for Neural Machine Translation
Lin, Pin-Jie
Chang, Ernie
Shi, Yangyang
Chandra, Vikas
Computation and Language
Machine Learning
Past vocabulary learning techniques identify relevant vocabulary before training, relying on statistical and entropy-based assumptions that largely neglect the role of model training. Empirically, we observe that trained translation models are induced to use a byte-pair encoding (BPE) vocabulary subset distinct from the original BPE vocabulary, leading to performance improvements when retrained with the induced vocabulary. In this paper, we analyze this discrepancy in neural machine translation by examining vocabulary and entropy shifts during self-training--where each iteration generates a labeled dataset by pairing source sentences with the model's predictions to define a new vocabulary. Building on these insights, we propose self-vocabularizing training, an iterative method that self-selects a smaller, more optimal vocabulary, yielding up to a 1.49 BLEU improvement. Moreover, we find that deeper model architectures lead to both an increase in unique token usage and a 6-8% reduction in vocabulary size.
title Self-Vocabularizing Training for Neural Machine Translation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2503.13837