Transformer tricks: Precomputing the first layer

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteur principal: Graef, Nils
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917611212636160
author Graef, Nils
author_facet Graef, Nils
contents This micro-paper describes a trick to speed up inference of transformers with RoPE (such as LLaMA, Mistral, PaLM, and Gemma). For these models, a large portion of the first transformer layer can be precomputed, which results in slightly lower latency and lower cost-per-token. Because this trick optimizes only one layer, the relative savings depend on the total number of layers. For example, the maximum savings for a model with only 4 layers (such as Whisper tiny) is limited to 25%, while a 32-layer model is limited to 3% savings. See https://github.com/OpenMachine-ai/transformer-tricks for code and more transformer tricks.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13388
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transformer tricks: Precomputing the first layer
Graef, Nils
Machine Learning
This micro-paper describes a trick to speed up inference of transformers with RoPE (such as LLaMA, Mistral, PaLM, and Gemma). For these models, a large portion of the first transformer layer can be precomputed, which results in slightly lower latency and lower cost-per-token. Because this trick optimizes only one layer, the relative savings depend on the total number of layers. For example, the maximum savings for a model with only 4 layers (such as Whisper tiny) is limited to 25%, while a 32-layer model is limited to 3% savings. See https://github.com/OpenMachine-ai/transformer-tricks for code and more transformer tricks.
title Transformer tricks: Precomputing the first layer
topic Machine Learning
url https://arxiv.org/abs/2402.13388