EpMAN: Episodic Memory AttentioN for Generalizing to Longer Contexts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chaudhury, Subhajit, Das, Payel, Swaminathan, Sarathkrishna, Kollias, Georgios, Nelson, Elliot, Pahwa, Khushbu, Pedapati, Tejaswini, Melnyk, Igor, Riemer, Matthew
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916622517665792
author Chaudhury, Subhajit
Das, Payel
Swaminathan, Sarathkrishna
Kollias, Georgios
Nelson, Elliot
Pahwa, Khushbu
Pedapati, Tejaswini
Melnyk, Igor
Riemer, Matthew
author_facet Chaudhury, Subhajit
Das, Payel
Swaminathan, Sarathkrishna
Kollias, Georgios
Nelson, Elliot
Pahwa, Khushbu
Pedapati, Tejaswini
Melnyk, Igor
Riemer, Matthew
contents Recent advances in Large Language Models (LLMs) have yielded impressive successes on many language tasks. However, efficient processing of long contexts using LLMs remains a significant challenge. We introduce \textbf{EpMAN} -- a method for processing long contexts in an \textit{episodic memory} module while \textit{holistically attending to} semantically relevant context chunks. The output of \textit{episodic attention} is then used to reweigh the decoder's self-attention to the stored KV cache of the context during training and generation. When an LLM decoder is trained using \textbf{EpMAN}, its performance on multiple challenging single-hop long-context recall and question-answering benchmarks is found to be stronger and more robust across the range from 16k to 256k tokens than baseline decoders trained with self-attention, and popular retrieval-augmented generation frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14280
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EpMAN: Episodic Memory AttentioN for Generalizing to Longer Contexts
Chaudhury, Subhajit
Das, Payel
Swaminathan, Sarathkrishna
Kollias, Georgios
Nelson, Elliot
Pahwa, Khushbu
Pedapati, Tejaswini
Melnyk, Igor
Riemer, Matthew
Computation and Language
Artificial Intelligence
Recent advances in Large Language Models (LLMs) have yielded impressive successes on many language tasks. However, efficient processing of long contexts using LLMs remains a significant challenge. We introduce \textbf{EpMAN} -- a method for processing long contexts in an \textit{episodic memory} module while \textit{holistically attending to} semantically relevant context chunks. The output of \textit{episodic attention} is then used to reweigh the decoder's self-attention to the stored KV cache of the context during training and generation. When an LLM decoder is trained using \textbf{EpMAN}, its performance on multiple challenging single-hop long-context recall and question-answering benchmarks is found to be stronger and more robust across the range from 16k to 256k tokens than baseline decoders trained with self-attention, and popular retrieval-augmented generation frameworks.
title EpMAN: Episodic Memory AttentioN for Generalizing to Longer Contexts
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.14280