SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Renxi, Mu, Honglin, Ma, Liqun, Lin, Lizhi, Feng, Yunlong, Baldwin, Timothy, Han, Xudong, Li, Haonan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915745039908864
author Wang, Renxi
Mu, Honglin
Ma, Liqun
Lin, Lizhi
Feng, Yunlong
Baldwin, Timothy
Han, Xudong
Li, Haonan
author_facet Wang, Renxi
Mu, Honglin
Ma, Liqun
Lin, Lizhi
Feng, Yunlong
Baldwin, Timothy
Han, Xudong
Li, Haonan
contents Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchmark designed to assess citation-grounded long-context reasoning in academic writing. SCALAR leverages academic papers and their citation structure to automatically generate high-quality ground-truth labels without human annotation. It features controllable difficulty levels and a dynamic updating mechanism that mitigates data contamination. The benchmark includes two tasks: a multiple-choice QA format and a cloze-style citation prediction. We evaluate a range of state-of-the-art LLMs and find that the multiple-choice task effectively distinguishes model capabilities. While human experts achieve over 90% accuracy, most models struggle. The cloze-style task is even more challenging, with no model exceeding 50% accuracy. SCALAR provides a domain-grounded, continuously updating framework for tracking progress in citation-based long-context understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
Wang, Renxi
Mu, Honglin
Ma, Liqun
Lin, Lizhi
Feng, Yunlong
Baldwin, Timothy
Han, Xudong
Li, Haonan
Computation and Language
Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchmark designed to assess citation-grounded long-context reasoning in academic writing. SCALAR leverages academic papers and their citation structure to automatically generate high-quality ground-truth labels without human annotation. It features controllable difficulty levels and a dynamic updating mechanism that mitigates data contamination. The benchmark includes two tasks: a multiple-choice QA format and a cloze-style citation prediction. We evaluate a range of state-of-the-art LLMs and find that the multiple-choice task effectively distinguishes model capabilities. While human experts achieve over 90% accuracy, most models struggle. The cloze-style task is even more challenging, with no model exceeding 50% accuracy. SCALAR provides a domain-grounded, continuously updating framework for tracking progress in citation-based long-context understanding.
title SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
topic Computation and Language
url https://arxiv.org/abs/2502.13753