Composition of Memory Experts for Diffusion World Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stapf, Sebastian, Huertos, Pablo Acuaviva, Davtyan, Aram, Favaro, Paolo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910234355695616
author Stapf, Sebastian
Huertos, Pablo Acuaviva
Davtyan, Aram
Favaro, Paolo
author_facet Stapf, Sebastian
Huertos, Pablo Acuaviva
Davtyan, Aram
Favaro, Paolo
contents World models aim to predict plausible futures consistent with past observations, a capability central to planning and decision-making in reinforcement learning. Yet, existing architectures face a fundamental memory trade-off: transformers preserve local detail but are bottlenecked by quadratic attention, while recurrent and state-space models scale more efficiently but compress history at the cost of fidelity. To overcome this trade-off, we suggest decoupling future-past consistency from any single architecture and instead leveraging a set of specialized experts. We introduce a diffusion-based framework that integrates heterogeneous memory models through a contrastive product-of-experts formulation. Our approach instantiates three complementary roles: a short-term memory expert that captures fine local dynamics, a long-term memory expert that stores episodic history in external diffusion weights via lightweight test-time finetuning, and a spatial long-term memory expert that enforces geometric and spatial coherence. This compositional design avoids mode collapse and scales to long contexts without incurring a quadratic cost. Across simulated and real-world benchmarks, our method improves temporal consistency, recall of past observations, and navigation performance, establishing a novel paradigm for building and operating memory-augmented diffusion world models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18813
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Composition of Memory Experts for Diffusion World Models
Stapf, Sebastian
Huertos, Pablo Acuaviva
Davtyan, Aram
Favaro, Paolo
Machine Learning
Artificial Intelligence
World models aim to predict plausible futures consistent with past observations, a capability central to planning and decision-making in reinforcement learning. Yet, existing architectures face a fundamental memory trade-off: transformers preserve local detail but are bottlenecked by quadratic attention, while recurrent and state-space models scale more efficiently but compress history at the cost of fidelity. To overcome this trade-off, we suggest decoupling future-past consistency from any single architecture and instead leveraging a set of specialized experts. We introduce a diffusion-based framework that integrates heterogeneous memory models through a contrastive product-of-experts formulation. Our approach instantiates three complementary roles: a short-term memory expert that captures fine local dynamics, a long-term memory expert that stores episodic history in external diffusion weights via lightweight test-time finetuning, and a spatial long-term memory expert that enforces geometric and spatial coherence. This compositional design avoids mode collapse and scales to long contexts without incurring a quadratic cost. Across simulated and real-world benchmarks, our method improves temporal consistency, recall of past observations, and navigation performance, establishing a novel paradigm for building and operating memory-augmented diffusion world models.
title Composition of Memory Experts for Diffusion World Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.18813