Byte BPE Tokenization as an Inverse string Homomorphism

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geng, Saibo, Gambhir, Sankalp, Wendler, Chris, West, Robert
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916507203665920
author Geng, Saibo
Gambhir, Sankalp
Wendler, Chris
West, Robert
author_facet Geng, Saibo
Gambhir, Sankalp
Wendler, Chris
West, Robert
contents Tokenization is an important preprocessing step in the training and inference of large language models (LLMs). While there has been extensive research on the expressive power of the neural achitectures used in LLMs, the impact of tokenization has not been well understood. In this work, we demonstrate that tokenization, irrespective of the algorithm used, acts as an inverse homomorphism between strings and tokens. This suggests that the character space of the source language and the token space of the tokenized language are homomorphic, preserving the structural properties of the source language. Additionally, we explore the concept of proper tokenization, which refers to an unambiguous tokenization returned from the tokenizer. Our analysis reveals that the expressiveness of neural architectures in recognizing context-free languages is not affected by tokenization.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03160
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Byte BPE Tokenization as an Inverse string Homomorphism
Geng, Saibo
Gambhir, Sankalp
Wendler, Chris
West, Robert
Computation and Language
Tokenization is an important preprocessing step in the training and inference of large language models (LLMs). While there has been extensive research on the expressive power of the neural achitectures used in LLMs, the impact of tokenization has not been well understood. In this work, we demonstrate that tokenization, irrespective of the algorithm used, acts as an inverse homomorphism between strings and tokens. This suggests that the character space of the source language and the token space of the tokenized language are homomorphic, preserving the structural properties of the source language. Additionally, we explore the concept of proper tokenization, which refers to an unambiguous tokenization returned from the tokenizer. Our analysis reveals that the expressiveness of neural architectures in recognizing context-free languages is not affected by tokenization.
title Byte BPE Tokenization as an Inverse string Homomorphism
topic Computation and Language
url https://arxiv.org/abs/2412.03160