RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Petty, Jackson, Hu, Michael Y., Wang, Wentao, Ravfogel, Shauli, Merrill, William, Linzen, Tal
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908996185620480
author Petty, Jackson
Hu, Michael Y.
Wang, Wentao
Ravfogel, Shauli
Merrill, William
Linzen, Tal
author_facet Petty, Jackson
Hu, Michael Y.
Wang, Wentao
Ravfogel, Shauli
Merrill, William
Linzen, Tal
contents Large language models (LLMs) are increasingly used to solve complex tasks where they must retrieve and compose many pieces of in-context information in long reasoning chains. For many real-world tasks it is hard to accurately gauge how model performance and strategy change as task complexity grows. To evaluate models' complex reasoning capability in a scalable and verifiable way, we introduce RELIC (Recognition of Languages In-Context), a framework that evaluates an LLM's ability to decide whether a given string belongs to the context-free language (CFL) generated by a grammar presented in-context. CFL recognition allows us to modulate the intrinsic complexity of the problem by varying grammar size and string length and translate this asymptotic complexity into predictions for ideal LLM performance. We find that even the most advanced reasoning models perform poorly on RELIC, not only failing to appropriately scale their inference compute to keep pace with task difficulty, but even reducing the number of reasoning tokens they use as task complexity increases. We find that these decreases in compute accompany changes in reasoning strategy, as models move from identifying and implementing algorithmic solutions to guessing. For models whose full completions go uninspected, this manifests as ``quiet quitting'' on hard tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05205
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
Petty, Jackson
Hu, Michael Y.
Wang, Wentao
Ravfogel, Shauli
Merrill, William
Linzen, Tal
Computation and Language
Large language models (LLMs) are increasingly used to solve complex tasks where they must retrieve and compose many pieces of in-context information in long reasoning chains. For many real-world tasks it is hard to accurately gauge how model performance and strategy change as task complexity grows. To evaluate models' complex reasoning capability in a scalable and verifiable way, we introduce RELIC (Recognition of Languages In-Context), a framework that evaluates an LLM's ability to decide whether a given string belongs to the context-free language (CFL) generated by a grammar presented in-context. CFL recognition allows us to modulate the intrinsic complexity of the problem by varying grammar size and string length and translate this asymptotic complexity into predictions for ideal LLM performance. We find that even the most advanced reasoning models perform poorly on RELIC, not only failing to appropriately scale their inference compute to keep pace with task difficulty, but even reducing the number of reasoning tokens they use as task complexity increases. We find that these decreases in compute accompany changes in reasoning strategy, as models move from identifying and implementing algorithmic solutions to guessing. For models whose full completions go uninspected, this manifests as ``quiet quitting'' on hard tasks.
title RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
topic Computation and Language
url https://arxiv.org/abs/2506.05205