Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911735947984896 |
|---|---|
| author | Merin, Adril Putra Anugraha, David Purwarianti, Ayu Winata, Genta Indra |
| author_facet | Merin, Adril Putra Anugraha, David Purwarianti, Ayu Winata, Genta Indra |
| contents | Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmarks evaluate agents within a single session, ignoring past actions, stated preferences, and prior decisions that agents must integrate to fulfill personalized user goals. We introduce Momento, a benchmark for persistent agentic task completion in multi-session service environments, requiring agents to take consequential, tool-mediated actions while resolving temporal dependencies and evolving user goals across sessions. Experimental results reveal that current agents fail primarily through misestimation of user state, treating prior session history as a reliable proxy for current context rather than stale information requiring re-validation, highlighting a substantial gap between current agent capabilities and realistic long-horizon human-agent interaction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2606_00832 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations Merin, Adril Putra Anugraha, David Purwarianti, Ayu Winata, Genta Indra Computation and Language Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmarks evaluate agents within a single session, ignoring past actions, stated preferences, and prior decisions that agents must integrate to fulfill personalized user goals. We introduce Momento, a benchmark for persistent agentic task completion in multi-session service environments, requiring agents to take consequential, tool-mediated actions while resolving temporal dependencies and evolving user goals across sessions. Experimental results reveal that current agents fail primarily through misestimation of user state, treating prior session history as a reliable proxy for current context rather than stale information requiring re-validation, highlighting a substantial gap between current agent capabilities and realistic long-horizon human-agent interaction. |
| title | Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2606.00832 |