GATEAU: Selecting Influential Samples for Long Context Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Si, Shuzheng, Zhao, Haozhe, Chen, Gang, Li, Yunshui, Luo, Kangyang, Lv, Chuancheng, An, Kaikai, Qi, Fanchao, Chang, Baobao, Sun, Maosong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908539033747456
author Si, Shuzheng
Zhao, Haozhe
Chen, Gang
Li, Yunshui
Luo, Kangyang
Lv, Chuancheng
An, Kaikai
Qi, Fanchao
Chang, Baobao
Sun, Maosong
author_facet Si, Shuzheng
Zhao, Haozhe
Chen, Gang
Li, Yunshui
Luo, Kangyang
Lv, Chuancheng
An, Kaikai
Qi, Fanchao
Chang, Baobao
Sun, Maosong
contents Aligning large language models to handle instructions with extremely long contexts has yet to be fully investigated. Previous studies have attempted to scale up the available data volume by synthesizing long instruction-following samples, as constructing such a dataset tends to be challenging for annotators. However, a lack of a well-defined strategy for ensuring data quality may introduce low-quality samples and restrict the model's performance. Thus, we propose GATEAU, a novel framework to address the unique challenge of long context alignment by identifying the influential samples enriched with long-range dependency relations. Specifically, GATEAU measures the long-range dependencies from two essential aspects: the difficulty of generating target responses due to the long-range dependencies, and the difficulty of understanding long inputs due to such dependencies. Comprehensive experiments indicate that GATEAU effectively identifies influential samples, and the model trained on these selected samples exhibits better instruction-following and long-context understanding capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15633
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GATEAU: Selecting Influential Samples for Long Context Alignment
Si, Shuzheng
Zhao, Haozhe
Chen, Gang
Li, Yunshui
Luo, Kangyang
Lv, Chuancheng
An, Kaikai
Qi, Fanchao
Chang, Baobao
Sun, Maosong
Computation and Language
Artificial Intelligence
Aligning large language models to handle instructions with extremely long contexts has yet to be fully investigated. Previous studies have attempted to scale up the available data volume by synthesizing long instruction-following samples, as constructing such a dataset tends to be challenging for annotators. However, a lack of a well-defined strategy for ensuring data quality may introduce low-quality samples and restrict the model's performance. Thus, we propose GATEAU, a novel framework to address the unique challenge of long context alignment by identifying the influential samples enriched with long-range dependency relations. Specifically, GATEAU measures the long-range dependencies from two essential aspects: the difficulty of generating target responses due to the long-range dependencies, and the difficulty of understanding long inputs due to such dependencies. Comprehensive experiments indicate that GATEAU effectively identifies influential samples, and the model trained on these selected samples exhibits better instruction-following and long-context understanding capabilities.
title GATEAU: Selecting Influential Samples for Long Context Alignment
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.15633