Which Pieces Does Unigram Tokenization Really Need?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Land, Sander, Pinter, Yuval
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918437427609600
author Land, Sander
Pinter, Yuval
author_facet Land, Sander
Pinter, Yuval
contents The Unigram tokenization algorithm offers a probabilistic alternative to the greedy heuristics of Byte-Pair Encoding. Despite its theoretical elegance, its implementation in practice is complex, limiting its adoption to the SentencePiece package and adapters thereof. We bridge this gap between theory and practice by providing a clear guide to implementation and parameter choices. We also identify a simpler algorithm that accepts slightly higher training loss in exchange for improved compression.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12641
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Which Pieces Does Unigram Tokenization Really Need?
Land, Sander
Pinter, Yuval
Computation and Language
68T50
The Unigram tokenization algorithm offers a probabilistic alternative to the greedy heuristics of Byte-Pair Encoding. Despite its theoretical elegance, its implementation in practice is complex, limiting its adoption to the SentencePiece package and adapters thereof. We bridge this gap between theory and practice by providing a clear guide to implementation and parameter choices. We also identify a simpler algorithm that accepts slightly higher training loss in exchange for improved compression.
title Which Pieces Does Unigram Tokenization Really Need?
topic Computation and Language
68T50
url https://arxiv.org/abs/2512.12641