The Ups and Downs of Large Language Model Inference with Vocabulary Trimming by Language Heuristics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bogoychev, Nikolay, Chen, Pinzhen, Haddow, Barry, Birch, Alexandra
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917652003291136
author Bogoychev, Nikolay
Chen, Pinzhen
Haddow, Barry
Birch, Alexandra
author_facet Bogoychev, Nikolay
Chen, Pinzhen
Haddow, Barry
Birch, Alexandra
contents Deploying large language models (LLMs) encounters challenges due to intensive computational and memory requirements. Our research examines vocabulary trimming (VT) inspired by restricting embedding entries to the language of interest to bolster time and memory efficiency. While such modifications have been proven effective in tasks like machine translation, tailoring them to LLMs demands specific modifications given the diverse nature of LLM applications. We apply two language heuristics to trim the full vocabulary - Unicode-based script filtering and corpus-based selection - to different LLM families and sizes. The methods are straightforward, interpretable, and easy to implement. It is found that VT reduces the memory usage of small models by nearly 50% and has an upper bound of 25% improvement in generation speed. Yet, we reveal the limitations of these methods in that they do not perform consistently well for each language with diminishing returns in larger models.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09709
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The Ups and Downs of Large Language Model Inference with Vocabulary Trimming by Language Heuristics
Bogoychev, Nikolay
Chen, Pinzhen
Haddow, Barry
Birch, Alexandra
Computation and Language
Deploying large language models (LLMs) encounters challenges due to intensive computational and memory requirements. Our research examines vocabulary trimming (VT) inspired by restricting embedding entries to the language of interest to bolster time and memory efficiency. While such modifications have been proven effective in tasks like machine translation, tailoring them to LLMs demands specific modifications given the diverse nature of LLM applications. We apply two language heuristics to trim the full vocabulary - Unicode-based script filtering and corpus-based selection - to different LLM families and sizes. The methods are straightforward, interpretable, and easy to implement. It is found that VT reduces the memory usage of small models by nearly 50% and has an upper bound of 25% improvement in generation speed. Yet, we reveal the limitations of these methods in that they do not perform consistently well for each language with diminishing returns in larger models.
title The Ups and Downs of Large Language Model Inference with Vocabulary Trimming by Language Heuristics
topic Computation and Language
url https://arxiv.org/abs/2311.09709