LSTM-MAS: A Long Short-Term Memory Inspired Multi-Agent System for Long-Context Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Yichen, Yuan, Jiakang, Tu, Chongjun, Ye, Peng, Chen, Tao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908981600976896
author Jiang, Yichen
Yuan, Jiakang
Tu, Chongjun
Ye, Peng
Chen, Tao
author_facet Jiang, Yichen
Yuan, Jiakang
Tu, Chongjun
Ye, Peng
Chen, Tao
contents Effectively processing long contexts remains a fundamental yet unsolved challenge for large language models (LLMs). Existing single-LLM-based methods primarily reduce the context window or optimize the attention mechanism, but they often encounter additional computational costs or constrained expanded context length. While multi-agent-based frameworks can mitigate these limitations, they remain susceptible to the accumulation of errors and the propagation of hallucinations. In this work, we draw inspiration from the Long Short-Term Memory (LSTM) architecture to design a Multi-Agent System called LSTM-MAS, emulating LSTM's hierarchical information flow and gated memory mechanisms for long-context understanding. Specifically, LSTM-MAS organizes agents in a chained architecture, where each node comprises a worker agent for segment-level comprehension, a filter agent for redundancy reduction, a judge agent for continuous error detection, and a manager agent for globally regulates information propagation and retention, analogous to LSTM and its input gate, forget gate, constant error carousel unit, and output gate. These novel designs enable controlled information transfer and selective long-term dependency modeling across textual segments, which can effectively avoid error accumulation and hallucination propagation. We conducted an extensive evaluation of our method. Compared with the previous best multi-agent approach, CoA, our model achieves improvements of 97.97%, 65.75%, 122.19%, 39.61% and 10.80% on Narrative QA, Qasper, HotpotQA, 2WikiMQA and MuSiQue, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11913
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LSTM-MAS: A Long Short-Term Memory Inspired Multi-Agent System for Long-Context Understanding
Jiang, Yichen
Yuan, Jiakang
Tu, Chongjun
Ye, Peng
Chen, Tao
Computation and Language
Artificial Intelligence
68T42
I.2.11
Effectively processing long contexts remains a fundamental yet unsolved challenge for large language models (LLMs). Existing single-LLM-based methods primarily reduce the context window or optimize the attention mechanism, but they often encounter additional computational costs or constrained expanded context length. While multi-agent-based frameworks can mitigate these limitations, they remain susceptible to the accumulation of errors and the propagation of hallucinations. In this work, we draw inspiration from the Long Short-Term Memory (LSTM) architecture to design a Multi-Agent System called LSTM-MAS, emulating LSTM's hierarchical information flow and gated memory mechanisms for long-context understanding. Specifically, LSTM-MAS organizes agents in a chained architecture, where each node comprises a worker agent for segment-level comprehension, a filter agent for redundancy reduction, a judge agent for continuous error detection, and a manager agent for globally regulates information propagation and retention, analogous to LSTM and its input gate, forget gate, constant error carousel unit, and output gate. These novel designs enable controlled information transfer and selective long-term dependency modeling across textual segments, which can effectively avoid error accumulation and hallucination propagation. We conducted an extensive evaluation of our method. Compared with the previous best multi-agent approach, CoA, our model achieves improvements of 97.97%, 65.75%, 122.19%, 39.61% and 10.80% on Narrative QA, Qasper, HotpotQA, 2WikiMQA and MuSiQue, respectively.
title LSTM-MAS: A Long Short-Term Memory Inspired Multi-Agent System for Long-Context Understanding
topic Computation and Language
Artificial Intelligence
68T42
I.2.11
url https://arxiv.org/abs/2601.11913