In Case You Missed It: ARC 'Challenge' Is Not That Challenging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Borchmann, Łukasz
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910760047738880
author Borchmann, Łukasz
author_facet Borchmann, Łukasz
contents ARC Challenge appears more difficult than ARC Easy for modern LLMs primarily due to an evaluation setup that prevents direct comparison of answer choices rather than inherent complexity. Although some researchers have quietly shifted to a more appropriate scheme over the last year, the implications of this change have yet to be widely acknowledged. We highlight this overlooked shift, show how similar evaluation practices falsely imply reasoning deficits in other benchmarks, and demonstrate that fairer methods dramatically reduce performance gaps (e.g. on SIQA) and even yield superhuman results (OpenBookQA). In doing so, we reveal how evaluation shapes perceived difficulty and offer guidelines to ensure that multiple-choice evaluations accurately reflect actual model capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17758
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle In Case You Missed It: ARC 'Challenge' Is Not That Challenging
Borchmann, Łukasz
Computation and Language
Artificial Intelligence
ARC Challenge appears more difficult than ARC Easy for modern LLMs primarily due to an evaluation setup that prevents direct comparison of answer choices rather than inherent complexity. Although some researchers have quietly shifted to a more appropriate scheme over the last year, the implications of this change have yet to be widely acknowledged. We highlight this overlooked shift, show how similar evaluation practices falsely imply reasoning deficits in other benchmarks, and demonstrate that fairer methods dramatically reduce performance gaps (e.g. on SIQA) and even yield superhuman results (OpenBookQA). In doing so, we reveal how evaluation shapes perceived difficulty and offer guidelines to ensure that multiple-choice evaluations accurately reflect actual model capabilities.
title In Case You Missed It: ARC 'Challenge' Is Not That Challenging
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.17758