MeMo: Towards Language Models with Associative Memory Mechanisms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zanzotto, Fabio Massimo, Ruzzetti, Elena Sofia, Xompero, Giancarlo A., Ranaldi, Leonardo, Venditti, Davide, Ranaldi, Federico, Giannone, Cristina, Favalli, Andrea, Romagnoli, Raniero
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914616883281920
author Zanzotto, Fabio Massimo
Ruzzetti, Elena Sofia
Xompero, Giancarlo A.
Ranaldi, Leonardo
Venditti, Davide
Ranaldi, Federico
Giannone, Cristina
Favalli, Andrea
Romagnoli, Raniero
author_facet Zanzotto, Fabio Massimo
Ruzzetti, Elena Sofia
Xompero, Giancarlo A.
Ranaldi, Leonardo
Venditti, Davide
Ranaldi, Federico
Giannone, Cristina
Favalli, Andrea
Romagnoli, Raniero
contents Memorization is a fundamental ability of Transformer-based Large Language Models, achieved through learning. In this paper, we propose a paradigm shift by designing an architecture to memorize text directly, bearing in mind the principle that memorization precedes learning. We introduce MeMo, a novel architecture for language modeling that explicitly memorizes sequences of tokens in layered associative memories. By design, MeMo offers transparency and the possibility of model editing, including forgetting texts. We experimented with the MeMo architecture, showing the memorization power of the one-layer and the multi-layer configurations.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MeMo: Towards Language Models with Associative Memory Mechanisms
Zanzotto, Fabio Massimo
Ruzzetti, Elena Sofia
Xompero, Giancarlo A.
Ranaldi, Leonardo
Venditti, Davide
Ranaldi, Federico
Giannone, Cristina
Favalli, Andrea
Romagnoli, Raniero
Computation and Language
Artificial Intelligence
I.2.7; I.2.6; I.2.4
Memorization is a fundamental ability of Transformer-based Large Language Models, achieved through learning. In this paper, we propose a paradigm shift by designing an architecture to memorize text directly, bearing in mind the principle that memorization precedes learning. We introduce MeMo, a novel architecture for language modeling that explicitly memorizes sequences of tokens in layered associative memories. By design, MeMo offers transparency and the possibility of model editing, including forgetting texts. We experimented with the MeMo architecture, showing the memorization power of the one-layer and the multi-layer configurations.
title MeMo: Towards Language Models with Associative Memory Mechanisms
topic Computation and Language
Artificial Intelligence
I.2.7; I.2.6; I.2.4
url https://arxiv.org/abs/2502.12851