Structured Packing in LLM Training Improves Long Context Utilization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Staniszewski, Konrad, Tworkowski, Szymon, Jaszczur, Sebastian, Zhao, Yu, Michalewski, Henryk, Kuciński, Łukasz, Miłoś, Piotr
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915175967227904
author Staniszewski, Konrad
Tworkowski, Szymon
Jaszczur, Sebastian
Zhao, Yu
Michalewski, Henryk
Kuciński, Łukasz
Miłoś, Piotr
author_facet Staniszewski, Konrad
Tworkowski, Szymon
Jaszczur, Sebastian
Zhao, Yu
Michalewski, Henryk
Kuciński, Łukasz
Miłoś, Piotr
contents Recent advancements in long-context large language models have attracted significant attention, yet their practical applications often suffer from suboptimal context utilization. This study investigates structuring training data to enhance semantic interdependence, demonstrating that this approach effectively improves context utilization. To this end, we introduce the Structured Packing for Long Context (SPLiCe) method, which utilizes retrieval to collate mutually relevant documents into long and coherent training examples. We validate SPLiCe empirically across models of varying sizes -- 3B, 7B, and 13B -- achieving improved performance in long-context tasks, such as Qasper and HotpotQA. Remarkably, even brief fine-tuning with SPLiCe is sufficient to realize these benefits. Additionally, SPLiCe effectively mitigates the lost-in-middle phenomenon often observed in large models. Our comprehensive analysis of SPLiCe explores its design choices and reveals intriguing transfer effects; for instance, training on programming code enhances performance on natural language tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2312_17296
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Structured Packing in LLM Training Improves Long Context Utilization
Staniszewski, Konrad
Tworkowski, Szymon
Jaszczur, Sebastian
Zhao, Yu
Michalewski, Henryk
Kuciński, Łukasz
Miłoś, Piotr
Computation and Language
Recent advancements in long-context large language models have attracted significant attention, yet their practical applications often suffer from suboptimal context utilization. This study investigates structuring training data to enhance semantic interdependence, demonstrating that this approach effectively improves context utilization. To this end, we introduce the Structured Packing for Long Context (SPLiCe) method, which utilizes retrieval to collate mutually relevant documents into long and coherent training examples. We validate SPLiCe empirically across models of varying sizes -- 3B, 7B, and 13B -- achieving improved performance in long-context tasks, such as Qasper and HotpotQA. Remarkably, even brief fine-tuning with SPLiCe is sufficient to realize these benefits. Additionally, SPLiCe effectively mitigates the lost-in-middle phenomenon often observed in large models. Our comprehensive analysis of SPLiCe explores its design choices and reveals intriguing transfer effects; for instance, training on programming code enhances performance on natural language tasks.
title Structured Packing in LLM Training Improves Long Context Utilization
topic Computation and Language
url https://arxiv.org/abs/2312.17296