Gradual Forgetting: Logarithmic Compression for Extending Transformer Context Windows

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dickson, Billy, Tiganj, Zoran
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914114655223808
author Dickson, Billy
Tiganj, Zoran
author_facet Dickson, Billy
Tiganj, Zoran
contents Most approaches to long-context processing increase the complexity of the transformer's internal architecture by integrating mechanisms such as recurrence or auxiliary memory modules. In this work, we introduce an alternative approach that modifies the input representation itself, rather than the transformer architecture. Inspired by cognitive models of human memory, our method applies a scale-invariant logarithmic compression to the input tokens. The resulting compressed representation is processed by a standard, unmodified transformer, preserving architectural simplicity. We evaluate this approach on the WikiText-103 and PG-19 language modeling benchmarks, showing a reduction in perplexity compared to uncompressed baselines. Moreover, performance improves consistently with longer compressed temporal contexts, showing that input-level logarithmic compression is a simple and effective way to extend a transformer's long-range memory.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22109
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gradual Forgetting: Logarithmic Compression for Extending Transformer Context Windows
Dickson, Billy
Tiganj, Zoran
Computation and Language
Artificial Intelligence
Most approaches to long-context processing increase the complexity of the transformer's internal architecture by integrating mechanisms such as recurrence or auxiliary memory modules. In this work, we introduce an alternative approach that modifies the input representation itself, rather than the transformer architecture. Inspired by cognitive models of human memory, our method applies a scale-invariant logarithmic compression to the input tokens. The resulting compressed representation is processed by a standard, unmodified transformer, preserving architectural simplicity. We evaluate this approach on the WikiText-103 and PG-19 language modeling benchmarks, showing a reduction in perplexity compared to uncompressed baselines. Moreover, performance improves consistently with longer compressed temporal contexts, showing that input-level logarithmic compression is a simple and effective way to extend a transformer's long-range memory.
title Gradual Forgetting: Logarithmic Compression for Extending Transformer Context Windows
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.22109