Layer-Specific Scaling of Positional Encodings for Superior Long-Context Modeling

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Zhenghua, Ding, Yiran, Lv, Changze, Xu, Zhibo, Li, Tianlong, Shi, Tianyuan, Zheng, Xiaoqing, Huang, Xuanjing
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916645221433344
author Wang, Zhenghua
Ding, Yiran
Lv, Changze
Xu, Zhibo
Li, Tianlong
Shi, Tianyuan
Zheng, Xiaoqing
Huang, Xuanjing
author_facet Wang, Zhenghua
Ding, Yiran
Lv, Changze
Xu, Zhibo
Li, Tianlong
Shi, Tianyuan
Zheng, Xiaoqing
Huang, Xuanjing
contents Although large language models (LLMs) have achieved significant progress in handling long-context inputs, they still suffer from the ``lost-in-the-middle'' problem, where crucial information in the middle of the context is often underrepresented or lost. Our extensive experiments reveal that this issue may arise from the rapid long-term decay in Rotary Position Embedding (RoPE). To address this problem, we propose a layer-specific positional encoding scaling method that assigns distinct scaling factors to each layer, slowing down the decay rate caused by RoPE to make the model pay more attention to the middle context. A specially designed genetic algorithm is employed to efficiently select the optimal scaling factors for each layer by incorporating Bezier curves to reduce the search space. Through comprehensive experimentation, we demonstrate that our method significantly alleviates the ``lost-in-the-middle'' problem. Our approach results in an average accuracy improvement of up to 20% on the Key-Value Retrieval dataset. Furthermore, we show that layer-specific interpolation, as opposed to uniform interpolation across all layers, enhances the model's extrapolation capabilities when combined with PI and Dynamic-NTK positional encoding schemes.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04355
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Layer-Specific Scaling of Positional Encodings for Superior Long-Context Modeling
Wang, Zhenghua
Ding, Yiran
Lv, Changze
Xu, Zhibo
Li, Tianlong
Shi, Tianyuan
Zheng, Xiaoqing
Huang, Xuanjing
Computation and Language
Although large language models (LLMs) have achieved significant progress in handling long-context inputs, they still suffer from the ``lost-in-the-middle'' problem, where crucial information in the middle of the context is often underrepresented or lost. Our extensive experiments reveal that this issue may arise from the rapid long-term decay in Rotary Position Embedding (RoPE). To address this problem, we propose a layer-specific positional encoding scaling method that assigns distinct scaling factors to each layer, slowing down the decay rate caused by RoPE to make the model pay more attention to the middle context. A specially designed genetic algorithm is employed to efficiently select the optimal scaling factors for each layer by incorporating Bezier curves to reduce the search space. Through comprehensive experimentation, we demonstrate that our method significantly alleviates the ``lost-in-the-middle'' problem. Our approach results in an average accuracy improvement of up to 20% on the Key-Value Retrieval dataset. Furthermore, we show that layer-specific interpolation, as opposed to uniform interpolation across all layers, enhances the model's extrapolation capabilities when combined with PI and Dynamic-NTK positional encoding schemes.
title Layer-Specific Scaling of Positional Encodings for Superior Long-Context Modeling
topic Computation and Language
url https://arxiv.org/abs/2503.04355