MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shen, Zhiyu, Wu, Ziming, Lai, Fuming, Lian, Shaobing, Rao, Yanghui
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908974709735424
author Shen, Zhiyu
Wu, Ziming
Lai, Fuming
Lian, Shaobing
Rao, Yanghui
author_facet Shen, Zhiyu
Wu, Ziming
Lai, Fuming
Lian, Shaobing
Rao, Yanghui
contents Maintaining consistency in long-term dialogues remains a fundamental challenge for LLMs, as standard retrieval mechanisms often fail to capture the temporal evolution of historical states. While memory-augmented frameworks offer a structured alternative, current systems rely on static prompting of closed-source models or suffer from ineffective training paradigms with sparse rewards. We introduce MemBuilder, a reinforcement learning framework that trains models to orchestrate multi-dimensional memory construction with attributed dense rewards. MemBuilder addresses two key challenges: (1) Sparse Trajectory-Level Rewards: we employ synthetic session-level question generation to provide dense intermediate rewards across extended trajectories; and (2) Multi-Dimensional Memory Attribution: we introduce contribution-aware gradient weighting that scales policy updates based on each component's downstream impact. Experimental results show that MemBuilder enables a 4B-parameter model to outperform state-of-the-art closed-source baselines, exhibiting strong generalization across long-term dialogue benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05488
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
Shen, Zhiyu
Wu, Ziming
Lai, Fuming
Lian, Shaobing
Rao, Yanghui
Computation and Language
Maintaining consistency in long-term dialogues remains a fundamental challenge for LLMs, as standard retrieval mechanisms often fail to capture the temporal evolution of historical states. While memory-augmented frameworks offer a structured alternative, current systems rely on static prompting of closed-source models or suffer from ineffective training paradigms with sparse rewards. We introduce MemBuilder, a reinforcement learning framework that trains models to orchestrate multi-dimensional memory construction with attributed dense rewards. MemBuilder addresses two key challenges: (1) Sparse Trajectory-Level Rewards: we employ synthetic session-level question generation to provide dense intermediate rewards across extended trajectories; and (2) Multi-Dimensional Memory Attribution: we introduce contribution-aware gradient weighting that scales policy updates based on each component's downstream impact. Experimental results show that MemBuilder enables a 4B-parameter model to outperform state-of-the-art closed-source baselines, exhibiting strong generalization across long-term dialogue benchmarks.
title MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
topic Computation and Language
url https://arxiv.org/abs/2601.05488