S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Encheng, Zhang, Jinouwen, Wu, Jianyu, Yu, Qiucheng, Tang, Chen, Li, Pengze, Wang, Lintao, Wang, Yizhou, Ma, Xinzhu, Tang, Shixiang, Wang, Aoran
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917540824875008
author Su, Encheng
Zhang, Jinouwen
Wu, Jianyu
Yu, Qiucheng
Tang, Chen
Li, Pengze
Wang, Lintao
Wang, Yizhou
Ma, Xinzhu
Tang, Shixiang
Wang, Aoran
author_facet Su, Encheng
Zhang, Jinouwen
Wu, Jianyu
Yu, Qiucheng
Tang, Chen
Li, Pengze
Wang, Lintao
Wang, Yizhou
Ma, Xinzhu
Tang, Shixiang
Wang, Aoran
contents Long-horizon interactive agents often accumulate large trajectory histories yet still fail to answer questions about earlier events reliably. We argue that the main bottleneck is not context length alone, but the trajectory-to-answer interface of long-term memory. When histories are stored as plain-text chunks and queried with standard retrieval-augmented generation (RAG), systems often retrieve locally relevant but chain-incomplete evidence, especially for spatial, temporal, repeated-event, and multi-hop state questions. We propose S3MEM, a structured scene-event episodic memory framework for long-horizon interactive question answering (QA). S3MEM writes trajectories into structured memory units, retrieves evidence through anchor-sensitive retrieval, and exposes a compact token-budget-aware evidence interface for answer-time inference. In this sense, S3MEM is a structured evidence harness that converts agent trajectories into query-aligned support. We evaluate S3MEM on two internal headline environments (Crafter, Jericho) and two out-of-family environments (SciWorld, ALFWorld). Under a shared frozen answer-time protocol, S3MEM consistently outperforms Vanilla RAG across all four environments, surpasses Graph-NoReader on Crafter, Jericho, and ALFWorld, and matches it on SciWorld while using dramatically fewer evidence tokens. Three adapted recent baselines -- A-MEM-inspired, MemoryOS-adapted, and LightMem-adapted -- improve over Vanilla RAG in several settings, but none matches S3MEM's overall accuracy-efficiency frontier. Overall, the evidence supports a bounded conclusion: under the current frozen answer-time protocol, structured writing and anchor-sensitive evidence routing provide a stronger accuracy-efficiency frontier for long-horizon interactive QA than more generic memory interfaces.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28831
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
Su, Encheng
Zhang, Jinouwen
Wu, Jianyu
Yu, Qiucheng
Tang, Chen
Li, Pengze
Wang, Lintao
Wang, Yizhou
Ma, Xinzhu
Tang, Shixiang
Wang, Aoran
Computation and Language
Artificial Intelligence
Long-horizon interactive agents often accumulate large trajectory histories yet still fail to answer questions about earlier events reliably. We argue that the main bottleneck is not context length alone, but the trajectory-to-answer interface of long-term memory. When histories are stored as plain-text chunks and queried with standard retrieval-augmented generation (RAG), systems often retrieve locally relevant but chain-incomplete evidence, especially for spatial, temporal, repeated-event, and multi-hop state questions. We propose S3MEM, a structured scene-event episodic memory framework for long-horizon interactive question answering (QA). S3MEM writes trajectories into structured memory units, retrieves evidence through anchor-sensitive retrieval, and exposes a compact token-budget-aware evidence interface for answer-time inference. In this sense, S3MEM is a structured evidence harness that converts agent trajectories into query-aligned support. We evaluate S3MEM on two internal headline environments (Crafter, Jericho) and two out-of-family environments (SciWorld, ALFWorld). Under a shared frozen answer-time protocol, S3MEM consistently outperforms Vanilla RAG across all four environments, surpasses Graph-NoReader on Crafter, Jericho, and ALFWorld, and matches it on SciWorld while using dramatically fewer evidence tokens. Three adapted recent baselines -- A-MEM-inspired, MemoryOS-adapted, and LightMem-adapted -- improve over Vanilla RAG in several settings, but none matches S3MEM's overall accuracy-efficiency frontier. Overall, the evidence supports a bounded conclusion: under the current frozen answer-time protocol, structured writing and anchor-sensitive evidence routing provide a stronger accuracy-efficiency frontier for long-horizon interactive QA than more generic memory interfaces.
title S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.28831