Can an LLM Induce a Graph? Investigating Memory Drift and Context Length

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yousuf, Raquib Bin, Khatri, Aadyant, Xu, Shengzhe, Sharma, Mandar, Ramakrishnan, Naren
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916988297674752
author Yousuf, Raquib Bin
Khatri, Aadyant
Xu, Shengzhe
Sharma, Mandar
Ramakrishnan, Naren
author_facet Yousuf, Raquib Bin
Khatri, Aadyant
Xu, Shengzhe
Sharma, Mandar
Ramakrishnan, Naren
contents Recently proposed evaluation benchmarks aim to characterize the effective context length and the forgetting tendencies of large language models (LLMs). However, these benchmarks often rely on simplistic 'needle in a haystack' retrieval or continuation tasks that may not accurately reflect the performance of these models in information-dense scenarios. Thus, rather than simple next token prediction, we argue for evaluating these models on more complex reasoning tasks that requires them to induce structured relational knowledge from the text - such as graphs from potentially noisy natural language content. While the input text can be viewed as generated in terms of a graph, its structure is not made explicit and connections must be induced from distributed textual cues, separated by long contexts and interspersed with irrelevant information. Our findings reveal that LLMs begin to exhibit memory drift and contextual forgetting at much shorter effective lengths when tasked with this form of relational reasoning, compared to what existing benchmarks suggest. With these findings, we offer recommendations for the optimal use of popular LLMs for complex reasoning tasks. We further show that even models specialized for reasoning, such as OpenAI o1, remain vulnerable to early memory drift in these settings. These results point to significant limitations in the models' ability to abstract structured knowledge from unstructured input and highlight the need for architectural adaptations to improve long-range reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03611
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
Yousuf, Raquib Bin
Khatri, Aadyant
Xu, Shengzhe
Sharma, Mandar
Ramakrishnan, Naren
Computation and Language
Artificial Intelligence
Machine Learning
Recently proposed evaluation benchmarks aim to characterize the effective context length and the forgetting tendencies of large language models (LLMs). However, these benchmarks often rely on simplistic 'needle in a haystack' retrieval or continuation tasks that may not accurately reflect the performance of these models in information-dense scenarios. Thus, rather than simple next token prediction, we argue for evaluating these models on more complex reasoning tasks that requires them to induce structured relational knowledge from the text - such as graphs from potentially noisy natural language content. While the input text can be viewed as generated in terms of a graph, its structure is not made explicit and connections must be induced from distributed textual cues, separated by long contexts and interspersed with irrelevant information. Our findings reveal that LLMs begin to exhibit memory drift and contextual forgetting at much shorter effective lengths when tasked with this form of relational reasoning, compared to what existing benchmarks suggest. With these findings, we offer recommendations for the optimal use of popular LLMs for complex reasoning tasks. We further show that even models specialized for reasoning, such as OpenAI o1, remain vulnerable to early memory drift in these settings. These results point to significant limitations in the models' ability to abstract structured knowledge from unstructured input and highlight the need for architectural adaptations to improve long-range reasoning.
title Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.03611