Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Tavakoli, Mohammad, Salemi, Alireza, Ye, Carrie, Abdalla, Mohamed, Zamani, Hamed, Mitchell, J Ross |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LaMP-QA: A Benchmark for Personalized Long-form Question Answering
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
Learning from Natural Language Feedback for Personalized Question Answering
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
Improving User Privacy in Personalized Generation: Client-Side Retrieval-Augmented Modification of Server-Side Generated Speculations
by: Salemi, Alireza, et al.
Published: (2026)
by: Salemi, Alireza, et al.
Published: (2026)
Learning to Reason for Multi-Step Retrieval of Personal Context in Personalized Question Answering
by: Amirizaniani, Maryam, et al.
Published: (2026)
by: Amirizaniani, Maryam, et al.
Published: (2026)
Towards a Search Engine for Machines: Unified Ranking for Multiple Retrieval-Augmented Large Language Models
by: Salemi, Alireza, et al.
Published: (2024)
by: Salemi, Alireza, et al.
Published: (2024)
Learning to Rank for Multiple Retrieval-Augmented Models through Iterative Utility Maximization
by: Salemi, Alireza, et al.
Published: (2024)
by: Salemi, Alireza, et al.
Published: (2024)
Evaluating Retrieval Quality in Retrieval-Augmented Generation
by: Salemi, Alireza, et al.
Published: (2024)
by: Salemi, Alireza, et al.
Published: (2024)
CIIR@LiveRAG 2025: Optimizing Multi-Agent Retrieval Augmented Generation through Self-Training
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
Plan-and-Refine: Diverse and Comprehensive Retrieval-Augmented Generation
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
Optimization Methods for Personalizing Large Language Models through Retrieval Augmentation
by: Salemi, Alireza, et al.
Published: (2024)
by: Salemi, Alireza, et al.
Published: (2024)
Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
Evaluation of Agents under Simulated AI Marketplace Dynamics
by: Kim, To Eun, et al.
Published: (2026)
by: Kim, To Eun, et al.
Published: (2026)
Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback
by: Alam, Md Zarif Ul, et al.
Published: (2026)
by: Alam, Md Zarif Ul, et al.
Published: (2026)
Retrieval-Enhanced Machine Learning: Synthesis and Opportunities
by: Kim, To Eun, et al.
Published: (2024)
by: Kim, To Eun, et al.
Published: (2024)
Benchmarking Information Retrieval Models on Complex Retrieval Tasks
by: Killingback, Julian, et al.
Published: (2025)
by: Killingback, Julian, et al.
Published: (2025)
Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question Answering
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
Accelerating Retrieval-Augmented Generation
by: Quinn, Derrick, et al.
Published: (2024)
by: Quinn, Derrick, et al.
Published: (2024)
Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
by: Jin, Bowen, et al.
Published: (2025)
by: Jin, Bowen, et al.
Published: (2025)
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
by: Tavakoli, Leila, et al.
Published: (2025)
by: Tavakoli, Leila, et al.
Published: (2025)
Distillation and Refinement of Reasoning in Small Language Models for Document Re-ranking
by: Samarinas, Chris, et al.
Published: (2025)
by: Samarinas, Chris, et al.
Published: (2025)
PerLTQA: A Personal Long-Term Memory Dataset for Memory Classification, Retrieval, and Synthesis in Question Answering
by: Du, Yiming, et al.
Published: (2024)
by: Du, Yiming, et al.
Published: (2024)
TableRAG: Million-Token Table Understanding with Language Models
by: Chen, Si-An, et al.
Published: (2024)
by: Chen, Si-An, et al.
Published: (2024)
CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search
by: Zeng, Hansi, et al.
Published: (2026)
by: Zeng, Hansi, et al.
Published: (2026)
Millions of $\text{GeAR}$-s: Extending GraphRAG to Millions of Documents
by: Shen, Zhili, et al.
Published: (2025)
by: Shen, Zhili, et al.
Published: (2025)
AI-Powered Assistant for Long-Term Access to RHIC Knowledge
by: Atif, Mohammad, et al.
Published: (2025)
by: Atif, Mohammad, et al.
Published: (2025)
APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI
by: Banerjee, Pratyay, et al.
Published: (2026)
by: Banerjee, Pratyay, et al.
Published: (2026)
Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous Decoding
by: Zeng, Hansi, et al.
Published: (2024)
by: Zeng, Hansi, et al.
Published: (2024)
Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization
by: Zamani, Hamed, et al.
Published: (2024)
by: Zamani, Hamed, et al.
Published: (2024)
Enhancing Long Context Performance in LLMs Through Inner Loop Query Mechanism
by: Tang, Yimin, et al.
Published: (2024)
by: Tang, Yimin, et al.
Published: (2024)
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
by: Chen, Yu, et al.
Published: (2026)
by: Chen, Yu, et al.
Published: (2026)
Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning
by: Samarinas, Chris, et al.
Published: (2026)
by: Samarinas, Chris, et al.
Published: (2026)
ProCIS: A Benchmark for Proactive Retrieval in Conversations
by: Samarinas, Chris, et al.
Published: (2024)
by: Samarinas, Chris, et al.
Published: (2024)
Structured Distillation for Personalized Agent Memory: 11x Token Reduction with Retrieval Preservation
by: Lewis, Sydney
Published: (2026)
by: Lewis, Sydney
Published: (2026)
Paths of A Million People: Extracting Life Trajectories from Wikipedia
by: Zhang, Ying, et al.
Published: (2024)
by: Zhang, Ying, et al.
Published: (2024)
Inferential Question Answering
by: Mozafari, Jamshid, et al.
Published: (2026)
by: Mozafari, Jamshid, et al.
Published: (2026)
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
by: Xing, Tiancheng, et al.
Published: (2025)
by: Xing, Tiancheng, et al.
Published: (2025)
Tuning LLMs by RAG Principles: Towards LLM-native Memory
by: Wei, Jiale, et al.
Published: (2025)
by: Wei, Jiale, et al.
Published: (2025)
Similar Items
-
LaMP-QA: A Benchmark for Personalized Long-form Question Answering
by: Salemi, Alireza, et al.
Published: (2025) -
Learning from Natural Language Feedback for Personalized Question Answering
by: Salemi, Alireza, et al.
Published: (2025) -
ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation
by: Salemi, Alireza, et al.
Published: (2025) -
Improving User Privacy in Personalized Generation: Client-Side Retrieval-Augmented Modification of Server-Side Generated Speculations
by: Salemi, Alireza, et al.
Published: (2026) -
Learning to Reason for Multi-Step Retrieval of Personal Context in Personalized Question Answering
by: Amirizaniani, Maryam, et al.
Published: (2026)