PHANTOM RECALL: When Familiar Puzzles Fool Smart Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mukhopadhyay, Souradeep, Baral, Rishabh, Mahajan, Nimeesh, Harish, Samhitha, RRV, Aswin, Parmar, Mihir, Nakamura, Mutsumi, Baral, Chitta
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915552825442304
author Mukhopadhyay, Souradeep
Baral, Rishabh
Mahajan, Nimeesh
Harish, Samhitha
RRV, Aswin
Parmar, Mihir
Nakamura, Mutsumi
Baral, Chitta
author_facet Mukhopadhyay, Souradeep
Baral, Rishabh
Mahajan, Nimeesh
Harish, Samhitha
RRV, Aswin
Parmar, Mihir
Nakamura, Mutsumi
Baral, Chitta
contents Large language models (LLMs) such as GPT, Gemini, and Claude often appear adept at solving classic logic puzzles--but how much genuine reasoning underlies their answers? Recent evidence suggests that these models frequently rely on memorized templates rather than reasoning from first principles. When puzzles are slightly modified, their performance collapses, revealing a striking fragility. In particular, we asked: Have LLMs addressed these issues? To what extent? How about perturbations to other puzzles? Is there a general way of reformulating the prompt so that the models do better? To examine these things systematically, we introduce PHANTOM RECALL, a benchmark comprising 25 well-known logic puzzles and 149 carefully designed perturbations that preserve reasoning structure but alter superficial details and solutions. We evaluate eleven leading LLMs and identify a recurring failure mode--phantom recall--where models confidently reproduce memorized solutions or spurious rationales that no longer fit the altered scenario. To probe and mitigate this issue, we contribute three tools: (i) an automated logical-equivalence judge to detect reasoning mismatches, (ii) a taxonomy of fine-grained reasoning error categories, and (iii) a prompting-based mitigation framework guided by these categories. Despite near-perfect accuracy on unmodified puzzles, models significantly underperform humans on perturbed ones, exhibiting both phantom recall and over-elaboration. Our findings reveal a crucial limitation: LLMs often fail to re-reason when contextual cues shift--highlighting the gap between linguistic fluency and logical understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11812
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
Mukhopadhyay, Souradeep
Baral, Rishabh
Mahajan, Nimeesh
Harish, Samhitha
RRV, Aswin
Parmar, Mihir
Nakamura, Mutsumi
Baral, Chitta
Computation and Language
Artificial Intelligence
Large language models (LLMs) such as GPT, Gemini, and Claude often appear adept at solving classic logic puzzles--but how much genuine reasoning underlies their answers? Recent evidence suggests that these models frequently rely on memorized templates rather than reasoning from first principles. When puzzles are slightly modified, their performance collapses, revealing a striking fragility. In particular, we asked: Have LLMs addressed these issues? To what extent? How about perturbations to other puzzles? Is there a general way of reformulating the prompt so that the models do better? To examine these things systematically, we introduce PHANTOM RECALL, a benchmark comprising 25 well-known logic puzzles and 149 carefully designed perturbations that preserve reasoning structure but alter superficial details and solutions. We evaluate eleven leading LLMs and identify a recurring failure mode--phantom recall--where models confidently reproduce memorized solutions or spurious rationales that no longer fit the altered scenario. To probe and mitigate this issue, we contribute three tools: (i) an automated logical-equivalence judge to detect reasoning mismatches, (ii) a taxonomy of fine-grained reasoning error categories, and (iii) a prompting-based mitigation framework guided by these categories. Despite near-perfect accuracy on unmodified puzzles, models significantly underperform humans on perturbed ones, exhibiting both phantom recall and over-elaboration. Our findings reveal a crucial limitation: LLMs often fail to re-reason when contextual cues shift--highlighting the gap between linguistic fluency and logical understanding.
title PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.11812