Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
Fuente:
arXiv
Saved in:
| Main Authors: | Merin, Adril Putra, Anugraha, David, Purwarianti, Ayu, Winata, Genta Indra |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
What Causes Knowledge Loss in Multilingual Language Models?
by: Khelli, Maria, et al.
Published: (2025)
by: Khelli, Maria, et al.
Published: (2025)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
by: Winata, Genta Indra, et al.
Published: (2026)
by: Winata, Genta Indra, et al.
Published: (2026)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
ProxyLM: Predicting Language Model Performance on Multilingual Tasks via Proxy Models
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
MINERS: Multilingual Language Models as Semantic Retrievers
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2026)
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2026)
T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
by: Chakraborty, Amartya, et al.
Published: (2025)
by: Chakraborty, Amartya, et al.
Published: (2025)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
by: Hudi, Frederikus, et al.
Published: (2025)
by: Hudi, Frederikus, et al.
Published: (2025)
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
by: Kuwanto, Garry, et al.
Published: (2024)
by: Kuwanto, Garry, et al.
Published: (2024)
Enhancing Natural Language Inference Performance with Knowledge Graph for COVID-19 Automated Fact-Checking in Indonesian Language
by: Muharram, Arief Purnama, et al.
Published: (2024)
by: Muharram, Arief Purnama, et al.
Published: (2024)
R3: Robust Rubric-Agnostic Reward Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
by: Elsetohy, Alaa, et al.
Published: (2026)
by: Elsetohy, Alaa, et al.
Published: (2026)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
by: He, Zexue, et al.
Published: (2026)
by: He, Zexue, et al.
Published: (2026)
Do Language Models Understand Honorific Systems in Javanese?
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2025)
Semantic Anchoring in Agentic Memory: Leveraging Linguistic Structures for Persistent Conversational Context
by: Chatterjee, Maitreyi, et al.
Published: (2025)
by: Chatterjee, Maitreyi, et al.
Published: (2025)
Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian Languages
by: Cahyawijaya, Samuel, et al.
Published: (2024)
by: Cahyawijaya, Samuel, et al.
Published: (2024)
Continual Learning in Machine Speech Chain Using Gradient Episodic Memory
by: Tyndall, Geoffrey, et al.
Published: (2024)
by: Tyndall, Geoffrey, et al.
Published: (2024)
Could We Have Had Better Multilingual LLMs If English Was Not the Central Language?
by: Diandaru, Ryandito, et al.
Published: (2024)
by: Diandaru, Ryandito, et al.
Published: (2024)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
by: Adila, Aulia, et al.
Published: (2024)
by: Adila, Aulia, et al.
Published: (2024)
Mixed-Session Conversation with Egocentric Memory
by: Jang, Jihyoung, et al.
Published: (2024)
by: Jang, Jihyoung, et al.
Published: (2024)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
by: Nur'aini, Khumaisa, et al.
Published: (2026)
by: Nur'aini, Khumaisa, et al.
Published: (2026)
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
Crosslingual Reasoning through Test-Time Scaling
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
DIAL-SUMMER: A Structured Evaluation Framework of Hierarchical Errors in Dialogue Summaries
by: Ramnath, Sahana, et al.
Published: (2026)
by: Ramnath, Sahana, et al.
Published: (2026)
Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability
by: Winata, Genta Indra, et al.
Published: (2025)
by: Winata, Genta Indra, et al.
Published: (2025)
Toward Multi-Session Personalized Conversation: A Large-Scale Dataset and Hierarchical Tree Framework for Implicit Reasoning
by: Li, Xintong, et al.
Published: (2025)
by: Li, Xintong, et al.
Published: (2025)
Language Surgery in Multilingual Large Language Models
by: Lopo, Joanito Agili, et al.
Published: (2025)
by: Lopo, Joanito Agili, et al.
Published: (2025)
DriveThru: a Document Extraction Platform and Benchmark Datasets for Indonesian Local Language Archives
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2024)
by: Farhansyah, Mohammad Rifqi, et al.
Published: (2024)
QLESS: A Quantized Approach for Data Valuation and Selection in Large Language Model Fine-Tuning
by: Ananta, Moses, et al.
Published: (2025)
by: Ananta, Moses, et al.
Published: (2025)
APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI
by: Banerjee, Pratyay, et al.
Published: (2026)
by: Banerjee, Pratyay, et al.
Published: (2026)
Similar Items
-
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025) -
What Causes Knowledge Loss in Multilingual Language Models?
by: Khelli, Maria, et al.
Published: (2025) -
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
by: Irawan, Patrick Amadeus, et al.
Published: (2024) -
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024) -
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
by: Winata, Genta Indra, et al.
Published: (2026)