Tokenisation via Convex Relaxations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tempus, Jan, Whittington, Philip, Schmidt, Craig W., Komm, Dennis, Pimentel, Tiago
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913153433993216
author Tempus, Jan
Whittington, Philip
Schmidt, Craig W.
Komm, Dennis
Pimentel, Tiago
author_facet Tempus, Jan
Whittington, Philip
Schmidt, Craig W.
Komm, Dennis
Pimentel, Tiago
contents Tokenisation is an integral part of the current NLP pipeline. Current tokenisation algorithms such as BPE and Unigram are greedy algorithms -- they make locally optimal decisions without considering the resulting vocabulary as a whole. We instead formulate tokeniser construction as a linear program and solve it using convex optimisation tools, yielding a new algorithm we call ConvexTok. We find ConvexTok consistently improves intrinsic tokenisation metrics and the bits-per-byte (BpB) achieved by language models; it also improves downstream task performance, but less consistently. Furthermore, ConvexTok allows the user to certify how far their tokeniser is from optimal, with respect to a certain objective, via a lower bound, and we empirically find it to be within 1\% of optimal at common vocabulary sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22821
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tokenisation via Convex Relaxations
Tempus, Jan
Whittington, Philip
Schmidt, Craig W.
Komm, Dennis
Pimentel, Tiago
Computation and Language
Machine Learning
Tokenisation is an integral part of the current NLP pipeline. Current tokenisation algorithms such as BPE and Unigram are greedy algorithms -- they make locally optimal decisions without considering the resulting vocabulary as a whole. We instead formulate tokeniser construction as a linear program and solve it using convex optimisation tools, yielding a new algorithm we call ConvexTok. We find ConvexTok consistently improves intrinsic tokenisation metrics and the bits-per-byte (BpB) achieved by language models; it also improves downstream task performance, but less consistently. Furthermore, ConvexTok allows the user to certify how far their tokeniser is from optimal, with respect to a certain objective, via a lower bound, and we empirically find it to be within 1\% of optimal at common vocabulary sizes.
title Tokenisation via Convex Relaxations
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.22821