LongReasonArena: A Long Reasoning Benchmark for Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ding, Jiayu, Ma, Shuming, Cui, Lei, Zheng, Nanning, Wei, Furu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915465799925760
author Ding, Jiayu
Ma, Shuming
Cui, Lei
Zheng, Nanning
Wei, Furu
author_facet Ding, Jiayu
Ma, Shuming
Cui, Lei
Zheng, Nanning
Wei, Furu
contents Existing long-context benchmarks for Large Language Models (LLMs) focus on evaluating comprehension of long inputs, while overlooking the evaluation of long reasoning abilities. To address this gap, we introduce LongReasonArena, a benchmark specifically designed to assess the long reasoning capabilities of LLMs. Our tasks require models to solve problems by executing multi-step algorithms that reflect key aspects of long reasoning, such as retrieval and backtracking. By controlling the inputs, the required reasoning length can be arbitrarily scaled, reaching up to 1 million tokens of reasoning for the most challenging tasks. Extensive evaluation results demonstrate that LongReasonArena presents a significant challenge for both open-source and proprietary LLMs. For instance, Deepseek-R1 achieves only 7.5% accuracy on our task. Further analysis also reveals that the accuracy exhibits a linear decline with respect to the logarithm of the expected number of reasoning steps. Our code and data is available at https://github.com/LongReasonArena/LongReasonArena.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19363
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LongReasonArena: A Long Reasoning Benchmark for Large Language Models
Ding, Jiayu
Ma, Shuming
Cui, Lei
Zheng, Nanning
Wei, Furu
Computation and Language
Artificial Intelligence
Existing long-context benchmarks for Large Language Models (LLMs) focus on evaluating comprehension of long inputs, while overlooking the evaluation of long reasoning abilities. To address this gap, we introduce LongReasonArena, a benchmark specifically designed to assess the long reasoning capabilities of LLMs. Our tasks require models to solve problems by executing multi-step algorithms that reflect key aspects of long reasoning, such as retrieval and backtracking. By controlling the inputs, the required reasoning length can be arbitrarily scaled, reaching up to 1 million tokens of reasoning for the most challenging tasks. Extensive evaluation results demonstrate that LongReasonArena presents a significant challenge for both open-source and proprietary LLMs. For instance, Deepseek-R1 achieves only 7.5% accuracy on our task. Further analysis also reveals that the accuracy exhibits a linear decline with respect to the logarithm of the expected number of reasoning steps. Our code and data is available at https://github.com/LongReasonArena/LongReasonArena.
title LongReasonArena: A Long Reasoning Benchmark for Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.19363