ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Datseris, Aleksis, Vassileva, Sylvia, Koychev, Ivan, Boytcheva, Svetla
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909822504402944
author Datseris, Aleksis
Vassileva, Sylvia
Koychev, Ivan
Boytcheva, Svetla
author_facet Datseris, Aleksis
Vassileva, Sylvia
Koychev, Ivan
Boytcheva, Svetla
contents This paper introduces a novel approach to position embeddings in transformer models, named "Exact Positional Embeddings" (ExPE). An absolute positional embedding method that can extrapolate to sequences of lengths longer than the ones it was trained on. Traditional transformer models rely on absolute or relative position embeddings to incorporate positional information into token embeddings, which often struggle with extrapolation to sequences longer than those seen during training. Our proposed method utilizes a novel embedding strategy that encodes exact positional information by overriding specific dimensions of the embedding vectors, thereby enabling a more precise representation of token positions. The proposed approach not only maintains the integrity of the original embeddings but also enhances the model's ability to generalize to more extended sequences. In causal language modeling, our ExPE embeddings significantly reduce perplexity compared to rotary and sinusoidal embeddings, when tested on sequences longer than those used in training.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19569
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
Datseris, Aleksis
Vassileva, Sylvia
Koychev, Ivan
Boytcheva, Svetla
Computation and Language
This paper introduces a novel approach to position embeddings in transformer models, named "Exact Positional Embeddings" (ExPE). An absolute positional embedding method that can extrapolate to sequences of lengths longer than the ones it was trained on. Traditional transformer models rely on absolute or relative position embeddings to incorporate positional information into token embeddings, which often struggle with extrapolation to sequences longer than those seen during training. Our proposed method utilizes a novel embedding strategy that encodes exact positional information by overriding specific dimensions of the embedding vectors, thereby enabling a more precise representation of token positions. The proposed approach not only maintains the integrity of the original embeddings but also enhances the model's ability to generalize to more extended sequences. In causal language modeling, our ExPE embeddings significantly reduce perplexity compared to rotary and sinusoidal embeddings, when tested on sequences longer than those used in training.
title ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
topic Computation and Language
url https://arxiv.org/abs/2509.19569