Fine-Tuning Transformers: Vocabulary Transfer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mosin, Vladislav, Samenko, Igor, Tikhonov, Alexey, Kozlovskii, Borislav, Yamshchikov, Ivan P.
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913217501986816
author Mosin, Vladislav
Samenko, Igor
Tikhonov, Alexey
Kozlovskii, Borislav
Yamshchikov, Ivan P.
author_facet Mosin, Vladislav
Samenko, Igor
Tikhonov, Alexey
Kozlovskii, Borislav
Yamshchikov, Ivan P.
contents Transformers are responsible for the vast majority of recent advances in natural language processing. The majority of practical natural language processing applications of these models are typically enabled through transfer learning. This paper studies if corpus-specific tokenization used for fine-tuning improves the resulting performance of the model. Through a series of experiments, we demonstrate that such tokenization combined with the initialization and fine-tuning strategy for the vocabulary tokens speeds up the transfer and boosts the performance of the fine-tuned model. We call this aspect of transfer facilitation vocabulary transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2112_14569
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Fine-Tuning Transformers: Vocabulary Transfer
Mosin, Vladislav
Samenko, Igor
Tikhonov, Alexey
Kozlovskii, Borislav
Yamshchikov, Ivan P.
Computation and Language
Artificial Intelligence
Machine Learning
68T50, 91F20
I.2.7
Transformers are responsible for the vast majority of recent advances in natural language processing. The majority of practical natural language processing applications of these models are typically enabled through transfer learning. This paper studies if corpus-specific tokenization used for fine-tuning improves the resulting performance of the model. Through a series of experiments, we demonstrate that such tokenization combined with the initialization and fine-tuning strategy for the vocabulary tokens speeds up the transfer and boosts the performance of the fine-tuned model. We call this aspect of transfer facilitation vocabulary transfer.
title Fine-Tuning Transformers: Vocabulary Transfer
topic Computation and Language
Artificial Intelligence
Machine Learning
68T50, 91F20
I.2.7
url https://arxiv.org/abs/2112.14569