Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Merin, Adril Putra, Anugraha, David, Purwarianti, Ayu, Winata, Genta Indra
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911735947984896
author Merin, Adril Putra
Anugraha, David
Purwarianti, Ayu
Winata, Genta Indra
author_facet Merin, Adril Putra
Anugraha, David
Purwarianti, Ayu
Winata, Genta Indra
contents Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmarks evaluate agents within a single session, ignoring past actions, stated preferences, and prior decisions that agents must integrate to fulfill personalized user goals. We introduce Momento, a benchmark for persistent agentic task completion in multi-session service environments, requiring agents to take consequential, tool-mediated actions while resolving temporal dependencies and evolving user goals across sessions. Experimental results reveal that current agents fail primarily through misestimation of user state, treating prior session history as a reliable proxy for current context rather than stale information requiring re-validation, highlighting a substantial gap between current agent capabilities and realistic long-horizon human-agent interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00832
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
Merin, Adril Putra
Anugraha, David
Purwarianti, Ayu
Winata, Genta Indra
Computation and Language
Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmarks evaluate agents within a single session, ignoring past actions, stated preferences, and prior decisions that agents must integrate to fulfill personalized user goals. We introduce Momento, a benchmark for persistent agentic task completion in multi-session service environments, requiring agents to take consequential, tool-mediated actions while resolving temporal dependencies and evolving user goals across sessions. Experimental results reveal that current agents fail primarily through misestimation of user state, treating prior session history as a reliable proxy for current context rather than stale information requiring re-validation, highlighting a substantial gap between current agent capabilities and realistic long-horizon human-agent interaction.
title Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
topic Computation and Language
url https://arxiv.org/abs/2606.00832