Out-of-Context Reasoning in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shaki, Jonathan, La Malfa, Emanuele, Wooldridge, Michael, Kraus, Sarit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915498609868800
author Shaki, Jonathan
La Malfa, Emanuele
Wooldridge, Michael
Kraus, Sarit
author_facet Shaki, Jonathan
La Malfa, Emanuele
Wooldridge, Michael
Kraus, Sarit
contents We study how large language models (LLMs) reason about memorized knowledge through simple binary relations such as equality ($=$), inequality ($<$), and inclusion ($\subset$). Unlike in-context reasoning, the axioms (e.g., $a < b, b < c$) are only seen during training and not provided in the task prompt (e.g., evaluating $a < c$). The tasks require one or more reasoning steps, and data aggregation from one or more sources, showing performance change with task complexity. We introduce a lightweight technique, out-of-context representation learning, which trains only new token embeddings on axioms and evaluates them on unseen tasks. Across reflexivity, symmetry, and transitivity tests, LLMs mostly perform statistically significant better than chance, making the correct answer extractable when testing multiple phrasing variations, but still fall short of consistent reasoning on every single query. Analysis shows that the learned embeddings are organized in structured ways, suggesting real relational understanding. Surprisingly, it also indicates that the core reasoning happens during the training, not inference.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10408
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Out-of-Context Reasoning in Large Language Models
Shaki, Jonathan
La Malfa, Emanuele
Wooldridge, Michael
Kraus, Sarit
Machine Learning
Computation and Language
We study how large language models (LLMs) reason about memorized knowledge through simple binary relations such as equality ($=$), inequality ($<$), and inclusion ($\subset$). Unlike in-context reasoning, the axioms (e.g., $a < b, b < c$) are only seen during training and not provided in the task prompt (e.g., evaluating $a < c$). The tasks require one or more reasoning steps, and data aggregation from one or more sources, showing performance change with task complexity. We introduce a lightweight technique, out-of-context representation learning, which trains only new token embeddings on axioms and evaluates them on unseen tasks. Across reflexivity, symmetry, and transitivity tests, LLMs mostly perform statistically significant better than chance, making the correct answer extractable when testing multiple phrasing variations, but still fall short of consistent reasoning on every single query. Analysis shows that the learned embeddings are organized in structured ways, suggesting real relational understanding. Surprisingly, it also indicates that the core reasoning happens during the training, not inference.
title Out-of-Context Reasoning in Large Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2503.10408