ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ho, Matthew, Si, Chen, Feng, Zhaoxiang, Yu, Fangxu, Yang, Yichi, Liu, Zhijian, Hu, Zhiting, Qin, Lianhui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914072864227328
author Ho, Matthew
Si, Chen
Feng, Zhaoxiang
Yu, Fangxu
Yang, Yichi
Liu, Zhijian
Hu, Zhiting
Qin, Lianhui
author_facet Ho, Matthew
Si, Chen
Feng, Zhaoxiang
Yu, Fangxu
Yang, Yichi
Liu, Zhijian
Hu, Zhiting
Qin, Lianhui
contents While inference-time scaling enables LLMs to carry out increasingly long and capable reasoning traces, the patterns and insights uncovered during these traces are immediately discarded once the context window is reset for a new query. External memory is a natural way to persist these discoveries, and recent work has shown clear benefits for reasoning-intensive tasks. We see an opportunity to make such memories more broadly reusable and scalable by moving beyond instance-based memory entries (e.g. exact query/response pairs, or summaries tightly coupled with the original problem context) toward concept-level memory: reusable, modular abstractions distilled from solution traces and stored in natural language. For future queries, relevant concepts are selectively retrieved and integrated into the prompt, enabling test-time continual learning without weight updates. Our design introduces new strategies for abstracting takeaways from rollouts and retrieving entries for new queries, promoting reuse and allowing memory to expand with additional experiences. We evaluate on ARC-AGI, a benchmark that stresses compositional generalization and abstract reasoning, making it a natural fit for concept memory. Our method yields a 7.5% relative gain over a strong no-memory baseline with performance continuing to scale with inference compute. We find abstract concepts to be the most consistent memory design, outscoring the baseline at all tested inference compute scales. Moreover, dynamically updating memory during test-time outperforms fixed settings, supporting the hypothesis that accumulating and abstracting patterns enables further solutions in a form of self-improvement. Code is available at https://github.com/matt-seb-ho/arc_memo.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory
Ho, Matthew
Si, Chen
Feng, Zhaoxiang
Yu, Fangxu
Yang, Yichi
Liu, Zhijian
Hu, Zhiting
Qin, Lianhui
Artificial Intelligence
Computation and Language
Machine Learning
While inference-time scaling enables LLMs to carry out increasingly long and capable reasoning traces, the patterns and insights uncovered during these traces are immediately discarded once the context window is reset for a new query. External memory is a natural way to persist these discoveries, and recent work has shown clear benefits for reasoning-intensive tasks. We see an opportunity to make such memories more broadly reusable and scalable by moving beyond instance-based memory entries (e.g. exact query/response pairs, or summaries tightly coupled with the original problem context) toward concept-level memory: reusable, modular abstractions distilled from solution traces and stored in natural language. For future queries, relevant concepts are selectively retrieved and integrated into the prompt, enabling test-time continual learning without weight updates. Our design introduces new strategies for abstracting takeaways from rollouts and retrieving entries for new queries, promoting reuse and allowing memory to expand with additional experiences. We evaluate on ARC-AGI, a benchmark that stresses compositional generalization and abstract reasoning, making it a natural fit for concept memory. Our method yields a 7.5% relative gain over a strong no-memory baseline with performance continuing to scale with inference compute. We find abstract concepts to be the most consistent memory design, outscoring the baseline at all tested inference compute scales. Moreover, dynamically updating memory during test-time outperforms fixed settings, supporting the hypothesis that accumulating and abstracting patterns enables further solutions in a form of self-improvement. Code is available at https://github.com/matt-seb-ho/arc_memo.
title ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.04439