From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Uddin, Md Nayem, Shubham, Kumar, Blanco, Eduardo, Baral, Chitta, Wang, Gengyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
by: Saeidi, Amir, et al.
Published: (2024)
by: Saeidi, Amir, et al.
Published: (2024)
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
by: RRV, Aswin, et al.
Published: (2024)
by: RRV, Aswin, et al.
Published: (2024)
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024)
by: Handa, Divij, et al.
Published: (2024)
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025)
by: RRV, Aswin, et al.
Published: (2025)
Asking and Answering Questions to Extract Event-Argument Structures
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
Generating Uncontextualized and Contextualized Questions for Document-Level Event Argument Extraction
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective
by: Rajput, Krishna Singh, et al.
Published: (2025)
by: Rajput, Krishna Singh, et al.
Published: (2025)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
by: Siingh, Shikhhar, et al.
Published: (2025)
by: Siingh, Shikhhar, et al.
Published: (2025)
MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
by: Zhang, Weizhi, et al.
Published: (2026)
by: Zhang, Weizhi, et al.
Published: (2026)
MemReader: From Passive to Active Extraction for Long-Term Agent Memory
by: Kang, Jingyi, et al.
Published: (2026)
by: Kang, Jingyi, et al.
Published: (2026)
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
by: Bei, Yuanchen, et al.
Published: (2026)
by: Bei, Yuanchen, et al.
Published: (2026)
Map&Make: Schema Guided Text to Table Generation
by: Ahuja, Naman, et al.
Published: (2025)
by: Ahuja, Naman, et al.
Published: (2025)
The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
by: Varshney, Neeraj, et al.
Published: (2023)
by: Varshney, Neeraj, et al.
Published: (2023)
Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents
by: Xu, Derong, et al.
Published: (2025)
by: Xu, Derong, et al.
Published: (2025)
ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
by: Chen, Tiantian, et al.
Published: (2026)
by: Chen, Tiantian, et al.
Published: (2026)
Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory
by: Sen, Sahil, et al.
Published: (2026)
by: Sen, Sahil, et al.
Published: (2026)
DimMem: Dimensional Structuring for Efficient Long-Term Agent Memory
by: Qiu, Wentao, et al.
Published: (2026)
by: Qiu, Wentao, et al.
Published: (2026)
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
MAPLE: Multi-Agent Adaptive Planning with Long-Term Memory for Table Reasoning
by: Bai, Ye, et al.
Published: (2025)
by: Bai, Ye, et al.
Published: (2025)
Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
by: Yu, Yi, et al.
Published: (2026)
by: Yu, Yi, et al.
Published: (2026)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
by: Fallah, Forouzan, et al.
Published: (2025)
by: Fallah, Forouzan, et al.
Published: (2025)
LLM-based Medical Assistant Personalization with Short- and Long-Term Memory Coordination
by: Zhang, Kai, et al.
Published: (2023)
by: Zhang, Kai, et al.
Published: (2023)
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
by: Ma, Junyu, et al.
Published: (2025)
by: Ma, Junyu, et al.
Published: (2025)
Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark
by: Gupta, Himanshu, et al.
Published: (2024)
by: Gupta, Himanshu, et al.
Published: (2024)
According to Me: Long-Term Personalized Referential Memory QA
by: Mei, Jingbiao, et al.
Published: (2026)
by: Mei, Jingbiao, et al.
Published: (2026)
Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization
by: Saeidi, Amir, et al.
Published: (2024)
by: Saeidi, Amir, et al.
Published: (2024)
MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents
by: Hu, Tianyu, et al.
Published: (2026)
by: Hu, Tianyu, et al.
Published: (2026)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
by: Patel, Maitreya, et al.
Published: (2023)
by: Patel, Maitreya, et al.
Published: (2023)
Mitigating Provenance-Role Collapse in Long-Term Agents via Typed Memory Representation
by: Jin, Zhengda, et al.
Published: (2026)
by: Jin, Zhengda, et al.
Published: (2026)
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
by: Tang, Zecheng, et al.
Published: (2026)
by: Tang, Zecheng, et al.
Published: (2026)
FadeMem: Biologically-Inspired Forgetting for Efficient Agent Memory
by: Wei, Lei, et al.
Published: (2026)
by: Wei, Lei, et al.
Published: (2026)
O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents
by: Wang, Piaohong, et al.
Published: (2025)
by: Wang, Piaohong, et al.
Published: (2025)
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
by: Handa, Divij, et al.
Published: (2025)
by: Handa, Divij, et al.
Published: (2025)
Preference-Aware Memory Update for Long-Term LLM Agents
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
by: Hu, Yulin, et al.
Published: (2026)
by: Hu, Yulin, et al.
Published: (2026)
M+: Extending MemoryLLM with Scalable Long-Term Memory
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Similar Items
-
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
by: Saeidi, Amir, et al.
Published: (2024) -
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
by: RRV, Aswin, et al.
Published: (2024) -
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
by: Uddin, Md Nayem, et al.
Published: (2024) -
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024) -
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025)