Long Input Benchmark for Russian Analysis

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Churin, Igor, Apishev, Murat, Tikhonova, Maria, Shevelev, Denis, Bulatov, Aydar, Kuratov, Yuri, Averkiev, Sergej, Fenogenova, Alena
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913458636718080
author Churin, Igor
Apishev, Murat
Tikhonova, Maria
Shevelev, Denis
Bulatov, Aydar
Kuratov, Yuri
Averkiev, Sergej
Fenogenova, Alena
author_facet Churin, Igor
Apishev, Murat
Tikhonova, Maria
Shevelev, Denis
Bulatov, Aydar
Kuratov, Yuri
Averkiev, Sergej
Fenogenova, Alena
contents Recent advancements in Natural Language Processing (NLP) have fostered the development of Large Language Models (LLMs) that can solve an immense variety of tasks. One of the key aspects of their application is their ability to work with long text documents and to process long sequences of tokens. This has created a demand for proper evaluation of long-context understanding. To address this need for the Russian language, we propose LIBRA (Long Input Benchmark for Russian Analysis), which comprises 21 adapted datasets to study the LLM's abilities to understand long texts thoroughly. The tests are divided into four complexity groups and allow the evaluation of models across various context lengths ranging from 4k up to 128k tokens. We provide the open-source datasets, codebase, and public leaderboard for LIBRA to guide forthcoming research.
format Preprint
id arxiv_https___arxiv_org_abs_2408_02439
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Long Input Benchmark for Russian Analysis
Churin, Igor
Apishev, Murat
Tikhonova, Maria
Shevelev, Denis
Bulatov, Aydar
Kuratov, Yuri
Averkiev, Sergej
Fenogenova, Alena
Computation and Language
Artificial Intelligence
Recent advancements in Natural Language Processing (NLP) have fostered the development of Large Language Models (LLMs) that can solve an immense variety of tasks. One of the key aspects of their application is their ability to work with long text documents and to process long sequences of tokens. This has created a demand for proper evaluation of long-context understanding. To address this need for the Russian language, we propose LIBRA (Long Input Benchmark for Russian Analysis), which comprises 21 adapted datasets to study the LLM's abilities to understand long texts thoroughly. The tests are divided into four complexity groups and allow the evaluation of models across various context lengths ranging from 4k up to 128k tokens. We provide the open-source datasets, codebase, and public leaderboard for LIBRA to guide forthcoming research.
title Long Input Benchmark for Russian Analysis
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2408.02439