Folded Context Condensation in Path Integral Formalism for Infinite Context Transformers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Paeng, Won-Gi, Kwon, Daesuk, Jeong, Kyungwon, Suh, Honggyo
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916715461345280
author Paeng, Won-Gi
Kwon, Daesuk
Jeong, Kyungwon
Suh, Honggyo
author_facet Paeng, Won-Gi
Kwon, Daesuk
Jeong, Kyungwon
Suh, Honggyo
contents In this work, we present a generalized formulation of the Transformer algorithm by reinterpreting its core mechanisms within the framework of Path Integral formalism. In this perspective, the attention mechanism is recast as a process that integrates all possible transition paths leading to future token states, with temporal evolution governed by the Feed-Forward Network. By systematically mapping each component of the Transformer to its counterpart in the Path Integral formulation, we obtain a more compact and efficient representation, in which the contextual information of a sequence is condensed into memory-like segments. These segments are recurrently processed across Transformer layers, enabling more effective long-term information retention. We validate the effectiveness of this approach through the Passkey retrieval task and a summarization task, demonstrating that the proposed method preserves historical information while exhibiting memory usage that scales linearly with sequence length. This contrasts with the non-linear memory growth typically observed in standard attention mechanisms. We expect that this quantum-inspired generalization of the Transformer architecture will open new avenues for enhancing both the efficiency and expressiveness of future Transformer models.
format Preprint
id arxiv_https___arxiv_org_abs_2405_04620
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Folded Context Condensation in Path Integral Formalism for Infinite Context Transformers
Paeng, Won-Gi
Kwon, Daesuk
Jeong, Kyungwon
Suh, Honggyo
High Energy Physics - Phenomenology
Artificial Intelligence
Computation and Language
Machine Learning
Neural and Evolutionary Computing
In this work, we present a generalized formulation of the Transformer algorithm by reinterpreting its core mechanisms within the framework of Path Integral formalism. In this perspective, the attention mechanism is recast as a process that integrates all possible transition paths leading to future token states, with temporal evolution governed by the Feed-Forward Network. By systematically mapping each component of the Transformer to its counterpart in the Path Integral formulation, we obtain a more compact and efficient representation, in which the contextual information of a sequence is condensed into memory-like segments. These segments are recurrently processed across Transformer layers, enabling more effective long-term information retention. We validate the effectiveness of this approach through the Passkey retrieval task and a summarization task, demonstrating that the proposed method preserves historical information while exhibiting memory usage that scales linearly with sequence length. This contrasts with the non-linear memory growth typically observed in standard attention mechanisms. We expect that this quantum-inspired generalization of the Transformer architecture will open new avenues for enhancing both the efficiency and expressiveness of future Transformer models.
title Folded Context Condensation in Path Integral Formalism for Infinite Context Transformers
topic High Energy Physics - Phenomenology
Artificial Intelligence
Computation and Language
Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2405.04620