Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yixiao, Zhu, Hanlin, Guo, Tianyu, Jiao, Jiantao, Sojoudi, Somayeh, Jordan, Michael I., Russell, Stuart, Mei, Song
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915767918788608
author Huang, Yixiao
Zhu, Hanlin
Guo, Tianyu
Jiao, Jiantao
Sojoudi, Somayeh
Jordan, Michael I.
Russell, Stuart
Mei, Song
author_facet Huang, Yixiao
Zhu, Hanlin
Guo, Tianyu
Jiao, Jiantao
Sojoudi, Somayeh
Jordan, Michael I.
Russell, Stuart
Mei, Song
contents Large language models (LLMs) can acquire new knowledge through fine-tuning, but this process exhibits a puzzling duality: models can generalize remarkably from new facts, yet are also prone to hallucinating incorrect information. However, the reasons for this phenomenon remain poorly understood. In this work, we argue that both behaviors stem from a single mechanism known as out-of-context reasoning (OCR): the ability to deduce implications by associating concepts, even those without a causal link. Our experiments across five prominent LLMs confirm that OCR indeed drives both generalization and hallucination, depending on whether the associated concepts are causally related. To build a rigorous theoretical understanding of this phenomenon, we then formalize OCR as a synthetic factual recall task. We empirically show that a one-layer single-head attention-only transformer with factorized output and value matrices can learn to solve this task, while a model with combined weights cannot, highlighting the crucial role of matrix factorization. Our theoretical analysis shows that the OCR capability can be attributed to the implicit bias of gradient descent, which favors solutions that minimize the nuclear norm of the combined output-value matrix. This mathematical structure explains why the model learns to associate facts and implications with high sample efficiency, regardless of whether the correlation is causal or merely spurious. Ultimately, our work provides a theoretical foundation for understanding the OCR phenomenon, offering a new lens for analyzing and mitigating undesirable behaviors from knowledge injection.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
Huang, Yixiao
Zhu, Hanlin
Guo, Tianyu
Jiao, Jiantao
Sojoudi, Somayeh
Jordan, Michael I.
Russell, Stuart
Mei, Song
Computation and Language
Machine Learning
Large language models (LLMs) can acquire new knowledge through fine-tuning, but this process exhibits a puzzling duality: models can generalize remarkably from new facts, yet are also prone to hallucinating incorrect information. However, the reasons for this phenomenon remain poorly understood. In this work, we argue that both behaviors stem from a single mechanism known as out-of-context reasoning (OCR): the ability to deduce implications by associating concepts, even those without a causal link. Our experiments across five prominent LLMs confirm that OCR indeed drives both generalization and hallucination, depending on whether the associated concepts are causally related. To build a rigorous theoretical understanding of this phenomenon, we then formalize OCR as a synthetic factual recall task. We empirically show that a one-layer single-head attention-only transformer with factorized output and value matrices can learn to solve this task, while a model with combined weights cannot, highlighting the crucial role of matrix factorization. Our theoretical analysis shows that the OCR capability can be attributed to the implicit bias of gradient descent, which favors solutions that minimize the nuclear norm of the combined output-value matrix. This mathematical structure explains why the model learns to associate facts and implications with high sample efficiency, regardless of whether the correlation is causal or merely spurious. Ultimately, our work provides a theoretical foundation for understanding the OCR phenomenon, offering a new lens for analyzing and mitigating undesirable behaviors from knowledge injection.
title Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.10887