Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Hawkins, John
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913826193014784
author Hawkins, John
author_facet Hawkins, John
contents Transformer-decoder language models are a core innovation in text based generative artificial intelligence. These models are being deployed as general-purpose intelligence systems in many applications. Central to their utility is the capacity to understand natural language commands and exploit the reasoning embedded in human text corpora to apply some form of reasoning process to a wide variety of novel tasks. To understand the limitations of this approach to generating reasoning we argue that we need to consider the architectural constraints of these systems. Consideration of the latent variable structure of transformer-decoder models allows us to design reasoning tasks that should probe the boundary of their capacity to reason. We present enigme, an open-source library for generating text-based puzzles to be used in training and evaluating reasoning skills within transformer-decoder models and future AI architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04914
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
Hawkins, John
Artificial Intelligence
Computation and Language
I.2.7
Transformer-decoder language models are a core innovation in text based generative artificial intelligence. These models are being deployed as general-purpose intelligence systems in many applications. Central to their utility is the capacity to understand natural language commands and exploit the reasoning embedded in human text corpora to apply some form of reasoning process to a wide variety of novel tasks. To understand the limitations of this approach to generating reasoning we argue that we need to consider the architectural constraints of these systems. Consideration of the latent variable structure of transformer-decoder models allows us to design reasoning tasks that should probe the boundary of their capacity to reason. We present enigme, an open-source library for generating text-based puzzles to be used in training and evaluating reasoning skills within transformer-decoder models and future AI architectures.
title Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models
topic Artificial Intelligence
Computation and Language
I.2.7
url https://arxiv.org/abs/2505.04914