From Tokens to Words: On the Inner Lexicon of LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kaplan, Guy, Oren, Matanel, Reif, Yuval, Schwartz, Roy
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915180168871936
author Kaplan, Guy
Oren, Matanel
Reif, Yuval
Schwartz, Roy
author_facet Kaplan, Guy
Oren, Matanel
Reif, Yuval
Schwartz, Roy
contents Natural language is composed of words, but modern large language models (LLMs) process sub-words as input. A natural question raised by this discrepancy is whether LLMs encode words internally, and if so how. We present evidence that LLMs engage in an intrinsic detokenization process, where sub-word sequences are combined into coherent whole-word representations at their last token. Our experiments show that this process primarily takes place within the early and middle layers of the model. We further demonstrate its robustness to arbitrary splits (e.g., "cats" to "ca" and "ts"), typos, and importantly-to out-of-vocabulary words: when feeding the last token internal representations of such words to the model as input, it can "understand" them as the complete word despite never seeing such representations as input during training. Our findings suggest that LLMs maintain a latent vocabulary beyond the tokenizer's scope. These insights provide a practical, finetuning-free application for expanding the vocabulary of pre-trained models. By enabling the addition of new vocabulary words, we reduce input length and inference iterations, which reduces both space and model latency, with little to no loss in model accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05864
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Tokens to Words: On the Inner Lexicon of LLMs
Kaplan, Guy
Oren, Matanel
Reif, Yuval
Schwartz, Roy
Computation and Language
Artificial Intelligence
Natural language is composed of words, but modern large language models (LLMs) process sub-words as input. A natural question raised by this discrepancy is whether LLMs encode words internally, and if so how. We present evidence that LLMs engage in an intrinsic detokenization process, where sub-word sequences are combined into coherent whole-word representations at their last token. Our experiments show that this process primarily takes place within the early and middle layers of the model. We further demonstrate its robustness to arbitrary splits (e.g., "cats" to "ca" and "ts"), typos, and importantly-to out-of-vocabulary words: when feeding the last token internal representations of such words to the model as input, it can "understand" them as the complete word despite never seeing such representations as input during training. Our findings suggest that LLMs maintain a latent vocabulary beyond the tokenizer's scope. These insights provide a practical, finetuning-free application for expanding the vocabulary of pre-trained models. By enabling the addition of new vocabulary words, we reduce input length and inference iterations, which reduces both space and model latency, with little to no loss in model accuracy.
title From Tokens to Words: On the Inner Lexicon of LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.05864