BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kuratov, Yuri, Bulatov, Aydar, Anokhin, Petr, Rodkin, Ivan, Sorokin, Dmitry, Sorokin, Artyom, Burtsev, Mikhail
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910686303485952
author Kuratov, Yuri
Bulatov, Aydar
Anokhin, Petr
Rodkin, Ivan
Sorokin, Dmitry
Sorokin, Artyom
Burtsev, Mikhail
author_facet Kuratov, Yuri
Bulatov, Aydar
Anokhin, Petr
Rodkin, Ivan
Sorokin, Dmitry
Sorokin, Artyom
Burtsev, Mikhail
contents In recent years, the input context sizes of large language models (LLMs) have increased dramatically. However, existing evaluation methods have not kept pace, failing to comprehensively assess the efficiency of models in handling long contexts. To bridge this gap, we introduce the BABILong benchmark, designed to test language models' ability to reason across facts distributed in extremely long documents. BABILong includes a diverse set of 20 reasoning tasks, including fact chaining, simple induction, deduction, counting, and handling lists/sets. These tasks are challenging on their own, and even more demanding when the required facts are scattered across long natural text. Our evaluations show that popular LLMs effectively utilize only 10-20\% of the context and their performance declines sharply with increased reasoning complexity. Among alternatives to in-context reasoning, Retrieval-Augmented Generation methods achieve a modest 60\% accuracy on single-fact question answering, independent of context length. Among context extension methods, the highest performance is demonstrated by recurrent memory transformers after fine-tuning, enabling the processing of lengths up to 50 million tokens. The BABILong benchmark is extendable to any length to support the evaluation of new upcoming models with increased capabilities, and we provide splits up to 10 million token lengths.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10149
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
Kuratov, Yuri
Bulatov, Aydar
Anokhin, Petr
Rodkin, Ivan
Sorokin, Dmitry
Sorokin, Artyom
Burtsev, Mikhail
Computation and Language
Artificial Intelligence
In recent years, the input context sizes of large language models (LLMs) have increased dramatically. However, existing evaluation methods have not kept pace, failing to comprehensively assess the efficiency of models in handling long contexts. To bridge this gap, we introduce the BABILong benchmark, designed to test language models' ability to reason across facts distributed in extremely long documents. BABILong includes a diverse set of 20 reasoning tasks, including fact chaining, simple induction, deduction, counting, and handling lists/sets. These tasks are challenging on their own, and even more demanding when the required facts are scattered across long natural text. Our evaluations show that popular LLMs effectively utilize only 10-20\% of the context and their performance declines sharply with increased reasoning complexity. Among alternatives to in-context reasoning, Retrieval-Augmented Generation methods achieve a modest 60\% accuracy on single-fact question answering, independent of context length. Among context extension methods, the highest performance is demonstrated by recurrent memory transformers after fine-tuning, enabling the processing of lengths up to 50 million tokens. The BABILong benchmark is extendable to any length to support the evaluation of new upcoming models with increased capabilities, and we provide splits up to 10 million token lengths.
title BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.10149