Length Generalization of Causal Transformers without Position Encoding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Jie, Ji, Tao, Wu, Yuanbin, Yan, Hang, Gui, Tao, Zhang, Qi, Huang, Xuanjing, Wang, Xiaoling
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917676932136960
author Wang, Jie
Ji, Tao
Wu, Yuanbin
Yan, Hang
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Wang, Xiaoling
author_facet Wang, Jie
Ji, Tao
Wu, Yuanbin
Yan, Hang
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Wang, Xiaoling
contents Generalizing to longer sentences is important for recent Transformer-based language models. Besides algorithms manipulating explicit position features, the success of Transformers without position encodings (NoPE) provides a new way to overcome the challenge. In this paper, we study the length generalization property of NoPE. We find that although NoPE can extend to longer sequences than the commonly used explicit position encodings, it still has a limited context length. We identify a connection between the failure of NoPE's generalization and the distraction of attention distributions. We propose a parameter-efficient tuning for searching attention heads' best temperature hyper-parameters, which substantially expands NoPE's context size. Experiments on long sequence language modeling, the synthetic passkey retrieval task and real-world long context tasks show that NoPE can achieve competitive performances with state-of-the-art length generalization algorithms. The source code is publicly accessible
format Preprint
id arxiv_https___arxiv_org_abs_2404_12224
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Length Generalization of Causal Transformers without Position Encoding
Wang, Jie
Ji, Tao
Wu, Yuanbin
Yan, Hang
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Wang, Xiaoling
Computation and Language
Generalizing to longer sentences is important for recent Transformer-based language models. Besides algorithms manipulating explicit position features, the success of Transformers without position encodings (NoPE) provides a new way to overcome the challenge. In this paper, we study the length generalization property of NoPE. We find that although NoPE can extend to longer sequences than the commonly used explicit position encodings, it still has a limited context length. We identify a connection between the failure of NoPE's generalization and the distraction of attention distributions. We propose a parameter-efficient tuning for searching attention heads' best temperature hyper-parameters, which substantially expands NoPE's context size. Experiments on long sequence language modeling, the synthetic passkey retrieval task and real-world long context tasks show that NoPE can achieve competitive performances with state-of-the-art length generalization algorithms. The source code is publicly accessible
title Length Generalization of Causal Transformers without Position Encoding
topic Computation and Language
url https://arxiv.org/abs/2404.12224