Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zheng, Yicong, McKee, Kevin L., Miconi, Thomas, Bugaud, Zacharie, van Gelderen, Mick, McCaleb, Jed
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908677538054144
author Zheng, Yicong
McKee, Kevin L.
Miconi, Thomas
Bugaud, Zacharie
van Gelderen, Mick
McCaleb, Jed
author_facet Zheng, Yicong
McKee, Kevin L.
Miconi, Thomas
Bugaud, Zacharie
van Gelderen, Mick
McCaleb, Jed
contents How to enable human-like long-term memory in large language models (LLMs) has been a central question for unlocking more general capabilities such as few-shot generalization. Existing memory frameworks and benchmarks focus on finding the optimal memory compression algorithm for higher performance in tasks that require recollection and sometimes further reasoning. However, such efforts have ended up building more human bias into the compression algorithm, through the search for the best prompts and memory architectures that suit specific benchmarks, rather than finding a general solution that would work on other data distributions. On the other hand, goal-directed search on uncompressed information could potentially exhibit superior performance because compression is lossy, and a predefined compression algorithm will not fit all raw data distributions. Here we present SUMER (Search in Uncompressed Memory via Experience Replay), an end-to-end reinforcement learning agent with verifiable reward (RLVR) that learns to use search tools to gather information and answer a target question. On the LoCoMo dataset for long-context conversation understanding, SUMER with Qwen2.5-7B-Instruct learned to use search tools and outperformed all other biased memory compression approaches and also the full-context baseline, reaching SOTA performance (43% gain over the prior best). We demonstrate that a simple search method applied to raw data outperforms goal-agnostic and biased compression algorithms in current long-context memory tasks, arguing for new paradigms and benchmarks that are more dynamic and autonomously scalable. Code for SUMER and all implemented baselines is publicly available at https://github.com/zycyc/SUMER.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
Zheng, Yicong
McKee, Kevin L.
Miconi, Thomas
Bugaud, Zacharie
van Gelderen, Mick
McCaleb, Jed
Computation and Language
Artificial Intelligence
Machine Learning
How to enable human-like long-term memory in large language models (LLMs) has been a central question for unlocking more general capabilities such as few-shot generalization. Existing memory frameworks and benchmarks focus on finding the optimal memory compression algorithm for higher performance in tasks that require recollection and sometimes further reasoning. However, such efforts have ended up building more human bias into the compression algorithm, through the search for the best prompts and memory architectures that suit specific benchmarks, rather than finding a general solution that would work on other data distributions. On the other hand, goal-directed search on uncompressed information could potentially exhibit superior performance because compression is lossy, and a predefined compression algorithm will not fit all raw data distributions. Here we present SUMER (Search in Uncompressed Memory via Experience Replay), an end-to-end reinforcement learning agent with verifiable reward (RLVR) that learns to use search tools to gather information and answer a target question. On the LoCoMo dataset for long-context conversation understanding, SUMER with Qwen2.5-7B-Instruct learned to use search tools and outperformed all other biased memory compression approaches and also the full-context baseline, reaching SOTA performance (43% gain over the prior best). We demonstrate that a simple search method applied to raw data outperforms goal-agnostic and biased compression algorithms in current long-context memory tasks, arguing for new paradigms and benchmarks that are more dynamic and autonomously scalable. Code for SUMER and all implemented baselines is publicly available at https://github.com/zycyc/SUMER.
title Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.21726