Evaluating Very Long-Term Conversational Memory of LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Maharana, Adyasha, Lee, Dong-Ho, Tulyakov, Sergey, Bansal, Mohit, Barbieri, Francesco, Fang, Yuwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation
by: Lee, Dong-Ho, et al.
Published: (2025)
by: Lee, Dong-Ho, et al.
Published: (2025)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
by: Maharana, Adyasha, et al.
Published: (2023)
by: Maharana, Adyasha, et al.
Published: (2023)
Adapt-$\infty$: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection
by: Maharana, Adyasha, et al.
Published: (2024)
by: Maharana, Adyasha, et al.
Published: (2024)
MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents
by: Hu, Tianyu, et al.
Published: (2026)
by: Hu, Tianyu, et al.
Published: (2026)
Soft Self-Consistency Improves Language Model Agents
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Preference-Aware Memory Update for Long-Term LLM Agents
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
by: Bei, Yuanchen, et al.
Published: (2026)
by: Bei, Yuanchen, et al.
Published: (2026)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
by: Xiao, Hanqi, et al.
Published: (2026)
by: Xiao, Hanqi, et al.
Published: (2026)
PRInTS: Reward Modeling for Long-Horizon Information Seeking
by: Lee, Jaewoo, et al.
Published: (2025)
by: Lee, Jaewoo, et al.
Published: (2025)
EngramaBench: Evaluating Long-Term Conversational Memory with Structured Graph Retrieval
by: Acuna, Julian
Published: (2026)
by: Acuna, Julian
Published: (2026)
Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
by: Zala, Abhay, et al.
Published: (2024)
by: Zala, Abhay, et al.
Published: (2024)
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
by: Yu, Hongli, et al.
Published: (2025)
by: Yu, Hongli, et al.
Published: (2025)
HyperMem: Hypergraph Memory for Long-Term Conversations
by: Yue, Juwei, et al.
Published: (2026)
by: Yue, Juwei, et al.
Published: (2026)
Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
by: Jiang, Yichen, et al.
Published: (2024)
by: Jiang, Yichen, et al.
Published: (2024)
A Human-Inspired Reading Agent with Gist Memory of Very Long Contexts
by: Lee, Kuang-Huei, et al.
Published: (2024)
by: Lee, Kuang-Huei, et al.
Published: (2024)
Effective Reasoning Chains Reduce Intrinsic Dimensionality
by: Prasad, Archiki, et al.
Published: (2026)
by: Prasad, Archiki, et al.
Published: (2026)
TrajAgent: An LLM-Agent Framework for Trajectory Modeling via Large-and-Small Model Collaboration
by: Du, Yuwei, et al.
Published: (2024)
by: Du, Yuwei, et al.
Published: (2024)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
by: Patil, Vaidehi, et al.
Published: (2025)
by: Patil, Vaidehi, et al.
Published: (2025)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
by: Saha, Swarnadeep, et al.
Published: (2023)
by: Saha, Swarnadeep, et al.
Published: (2023)
RecMem: Recurrence-based Memory Consolidation for Efficient and Effective Long-Running LLM Agents
by: Dai, Zijie, et al.
Published: (2026)
by: Dai, Zijie, et al.
Published: (2026)
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models
by: Ha, Hyeonjeong, et al.
Published: (2026)
by: Ha, Hyeonjeong, et al.
Published: (2026)
GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational Memory
by: Sun, Yushi, et al.
Published: (2026)
by: Sun, Yushi, et al.
Published: (2026)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
by: Chen, Justin Chih-Yao, et al.
Published: (2023)
by: Chen, Justin Chih-Yao, et al.
Published: (2023)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
by: Hase, Peter, et al.
Published: (2024)
by: Hase, Peter, et al.
Published: (2024)
Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
by: Bouchard, Dylan, et al.
Published: (2026)
by: Bouchard, Dylan, et al.
Published: (2026)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
by: Zala, Abhay, et al.
Published: (2023)
by: Zala, Abhay, et al.
Published: (2023)
ReGAL: Refactoring Programs to Discover Generalizable Abstractions
by: Stengel-Eskin, Elias, et al.
Published: (2024)
by: Stengel-Eskin, Elias, et al.
Published: (2024)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
by: Lin, Han, et al.
Published: (2023)
by: Lin, Han, et al.
Published: (2023)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
by: Deshpande, Darshan, et al.
Published: (2025)
by: Deshpande, Darshan, et al.
Published: (2025)
Multi-Attribute Steering of Language Models via Targeted Intervention
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024)
by: Dong, Zhichen, et al.
Published: (2024)
PSYCHE: A Multi-faceted Patient Simulation Framework for Evaluation of Psychiatric Assessment Conversational Agents
by: Lee, Jingoo, et al.
Published: (2025)
by: Lee, Jingoo, et al.
Published: (2025)
On the Structural Memory of LLM Agents
by: Zeng, Ruihong, et al.
Published: (2024)
by: Zeng, Ruihong, et al.
Published: (2024)
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
by: Levi, Elad, et al.
Published: (2025)
by: Levi, Elad, et al.
Published: (2025)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
by: Xiao, Hanqi, et al.
Published: (2025)
by: Xiao, Hanqi, et al.
Published: (2025)
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
Similar Items
-
REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation
by: Lee, Dong-Ho, et al.
Published: (2025) -
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
by: Maharana, Adyasha, et al.
Published: (2023) -
Adapt-$\infty$: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection
by: Maharana, Adyasha, et al.
Published: (2024) -
MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents
by: Hu, Tianyu, et al.
Published: (2026) -
Soft Self-Consistency Improves Language Model Agents
by: Wang, Han, et al.
Published: (2024)