SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915745039908864 |
|---|---|
| author | Wang, Renxi Mu, Honglin Ma, Liqun Lin, Lizhi Feng, Yunlong Baldwin, Timothy Han, Xudong Li, Haonan |
| author_facet | Wang, Renxi Mu, Honglin Ma, Liqun Lin, Lizhi Feng, Yunlong Baldwin, Timothy Han, Xudong Li, Haonan |
| contents | Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchmark designed to assess citation-grounded long-context reasoning in academic writing. SCALAR leverages academic papers and their citation structure to automatically generate high-quality ground-truth labels without human annotation. It features controllable difficulty levels and a dynamic updating mechanism that mitigates data contamination. The benchmark includes two tasks: a multiple-choice QA format and a cloze-style citation prediction. We evaluate a range of state-of-the-art LLMs and find that the multiple-choice task effectively distinguishes model capabilities. While human experts achieve over 90% accuracy, most models struggle. The cloze-style task is even more challenging, with no model exceeding 50% accuracy. SCALAR provides a domain-grounded, continuously updating framework for tracking progress in citation-based long-context understanding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_13753 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning Wang, Renxi Mu, Honglin Ma, Liqun Lin, Lizhi Feng, Yunlong Baldwin, Timothy Han, Xudong Li, Haonan Computation and Language Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchmark designed to assess citation-grounded long-context reasoning in academic writing. SCALAR leverages academic papers and their citation structure to automatically generate high-quality ground-truth labels without human annotation. It features controllable difficulty levels and a dynamic updating mechanism that mitigates data contamination. The benchmark includes two tasks: a multiple-choice QA format and a cloze-style citation prediction. We evaluate a range of state-of-the-art LLMs and find that the multiple-choice task effectively distinguishes model capabilities. While human experts achieve over 90% accuracy, most models struggle. The cloze-style task is even more challenging, with no model exceeding 50% accuracy. SCALAR provides a domain-grounded, continuously updating framework for tracking progress in citation-based long-context understanding. |
| title | SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2502.13753 |