Long Context Alignment with Short Instructions and Synthesized Positions

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Wenhao, Wang, Yizhong, Fu, Yao, Yue, Xiang, Zhu, Dawei, Li, Sujian
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913343240929280
author Wu, Wenhao
Wang, Yizhong
Fu, Yao
Yue, Xiang
Zhu, Dawei
Li, Sujian
author_facet Wu, Wenhao
Wang, Yizhong
Fu, Yao
Yue, Xiang
Zhu, Dawei
Li, Sujian
contents Effectively handling instructions with extremely long context remains a challenge for Large Language Models (LLMs), typically necessitating high-quality long data and substantial computational resources. This paper introduces Step-Skipping Alignment (SkipAlign), a new technique designed to enhance the long-context capabilities of LLMs in the phase of alignment without the need for additional efforts beyond training with original data length. SkipAlign is developed on the premise that long-range dependencies are fundamental to enhancing an LLM's capacity of long context. Departing from merely expanding the length of input samples, SkipAlign synthesizes long-range dependencies from the aspect of positions indices. This is achieved by the strategic insertion of skipped positions within instruction-following samples, which utilizes the semantic structure of the data to effectively expand the context. Through extensive experiments on base models with a variety of context window sizes, SkipAlign demonstrates its effectiveness across a spectrum of long-context tasks. Particularly noteworthy is that with a careful selection of the base model and alignment datasets, SkipAlign with only 6B parameters achieves it's best performance and comparable with strong baselines like GPT-3.5-Turbo-16K on LongBench.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Long Context Alignment with Short Instructions and Synthesized Positions
Wu, Wenhao
Wang, Yizhong
Fu, Yao
Yue, Xiang
Zhu, Dawei
Li, Sujian
Computation and Language
Effectively handling instructions with extremely long context remains a challenge for Large Language Models (LLMs), typically necessitating high-quality long data and substantial computational resources. This paper introduces Step-Skipping Alignment (SkipAlign), a new technique designed to enhance the long-context capabilities of LLMs in the phase of alignment without the need for additional efforts beyond training with original data length. SkipAlign is developed on the premise that long-range dependencies are fundamental to enhancing an LLM's capacity of long context. Departing from merely expanding the length of input samples, SkipAlign synthesizes long-range dependencies from the aspect of positions indices. This is achieved by the strategic insertion of skipped positions within instruction-following samples, which utilizes the semantic structure of the data to effectively expand the context. Through extensive experiments on base models with a variety of context window sizes, SkipAlign demonstrates its effectiveness across a spectrum of long-context tasks. Particularly noteworthy is that with a careful selection of the base model and alignment datasets, SkipAlign with only 6B parameters achieves it's best performance and comparable with strong baselines like GPT-3.5-Turbo-16K on LongBench.
title Long Context Alignment with Short Instructions and Synthesized Positions
topic Computation and Language
url https://arxiv.org/abs/2405.03939