BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chizhov, Pavel, Arnett, Catherine, Korotkova, Elizaveta, Yamshchikov, Ivan P.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910593257046016
author Chizhov, Pavel
Arnett, Catherine
Korotkova, Elizaveta
Yamshchikov, Ivan P.
author_facet Chizhov, Pavel
Arnett, Catherine
Korotkova, Elizaveta
Yamshchikov, Ivan P.
contents Language models can largely benefit from efficient tokenization. However, they still mostly utilize the classical BPE algorithm, a simple and reliable method. This has been shown to cause such issues as under-trained tokens and sub-optimal compression that may affect the downstream performance. We introduce Picky BPE, a modified BPE algorithm that carries out vocabulary refinement during tokenizer training. Our method improves vocabulary efficiency, eliminates under-trained tokens, and does not compromise text compression. Our experiments show that our method does not reduce the downstream performance, and in several cases improves it.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04599
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
Chizhov, Pavel
Arnett, Catherine
Korotkova, Elizaveta
Yamshchikov, Ivan P.
Computation and Language
Language models can largely benefit from efficient tokenization. However, they still mostly utilize the classical BPE algorithm, a simple and reliable method. This has been shown to cause such issues as under-trained tokens and sub-optimal compression that may affect the downstream performance. We introduce Picky BPE, a modified BPE algorithm that carries out vocabulary refinement during tokenizer training. Our method improves vocabulary efficiency, eliminates under-trained tokens, and does not compromise text compression. Our experiments show that our method does not reduce the downstream performance, and in several cases improves it.
title BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
topic Computation and Language
url https://arxiv.org/abs/2409.04599