Scaling Laws of RoPE-based Extrapolation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Xiaoran, Yan, Hang, Zhang, Shuo, An, Chenxin, Qiu, Xipeng, Lin, Dahua
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913263788228608
author Liu, Xiaoran
Yan, Hang
Zhang, Shuo
An, Chenxin
Qiu, Xipeng
Lin, Dahua
author_facet Liu, Xiaoran
Yan, Hang
Zhang, Shuo
An, Chenxin
Qiu, Xipeng
Lin, Dahua
contents The extrapolation capability of Large Language Models (LLMs) based on Rotary Position Embedding is currently a topic of considerable interest. The mainstream approach to addressing extrapolation with LLMs involves modifying RoPE by replacing 10000, the rotary base of $θ_n={10000}^{-2n/d}$ in the original RoPE, with a larger value and providing longer fine-tuning text. In this work, we first observe that fine-tuning a RoPE-based LLM with either a smaller or larger base in pre-training context length could significantly enhance its extrapolation performance. After that, we propose \textbf{\textit{Scaling Laws of RoPE-based Extrapolation}}, a unified framework from the periodic perspective, to describe the relationship between the extrapolation performance and base value as well as tuning context length. In this process, we also explain the origin of the RoPE-based extrapolation issue by \textbf{\textit{critical dimension for extrapolation}}. Besides these observations and analyses, we achieve extrapolation up to 1 million context length within only 16K training length on LLaMA2 7B and 13B.
format Preprint
id arxiv_https___arxiv_org_abs_2310_05209
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Scaling Laws of RoPE-based Extrapolation
Liu, Xiaoran
Yan, Hang
Zhang, Shuo
An, Chenxin
Qiu, Xipeng
Lin, Dahua
Computation and Language
Artificial Intelligence
The extrapolation capability of Large Language Models (LLMs) based on Rotary Position Embedding is currently a topic of considerable interest. The mainstream approach to addressing extrapolation with LLMs involves modifying RoPE by replacing 10000, the rotary base of $θ_n={10000}^{-2n/d}$ in the original RoPE, with a larger value and providing longer fine-tuning text. In this work, we first observe that fine-tuning a RoPE-based LLM with either a smaller or larger base in pre-training context length could significantly enhance its extrapolation performance. After that, we propose \textbf{\textit{Scaling Laws of RoPE-based Extrapolation}}, a unified framework from the periodic perspective, to describe the relationship between the extrapolation performance and base value as well as tuning context length. In this process, we also explain the origin of the RoPE-based extrapolation issue by \textbf{\textit{critical dimension for extrapolation}}. Besides these observations and analyses, we achieve extrapolation up to 1 million context length within only 16K training length on LLaMA2 7B and 13B.
title Scaling Laws of RoPE-based Extrapolation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2310.05209