Studying the Soupability of Documents in State Space Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jafari, Yasaman, Wang, Zixian, Bergen, Leon, Berg-Kirkpatrick, Taylor
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911407889448960
author Jafari, Yasaman
Wang, Zixian
Bergen, Leon
Berg-Kirkpatrick, Taylor
author_facet Jafari, Yasaman
Wang, Zixian
Bergen, Leon
Berg-Kirkpatrick, Taylor
contents We investigate whether hidden states from Structured State Space Models (SSMs) can be merged post hoc to support downstream reasoning. Inspired by model souping, we study document souping, a strategy where documents are encoded independently, and their representations are pooled, via simple operations like averaging, into a single context state. This approach enables modular encoding and reuse without reprocessing the full input for each query. We demonstrate that finetuned Mamba2 models with souped representations achieve competitive or superior performance across multi-hop QA, sparse retrieval, and long-document reasoning tasks compared to the standard monolithic encoding approach. For example, on the RACE and QuALITY benchmarks for long document question answering, this method substantially outperforms a traditional concatenation approach. Crucially, this modular design scales to hundreds of documents while delivering substantial savings in inference cost, unlocking new possibilities for large-scale corpus reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24033
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Studying the Soupability of Documents in State Space Models
Jafari, Yasaman
Wang, Zixian
Bergen, Leon
Berg-Kirkpatrick, Taylor
Computation and Language
Computational Engineering, Finance, and Science
Machine Learning
We investigate whether hidden states from Structured State Space Models (SSMs) can be merged post hoc to support downstream reasoning. Inspired by model souping, we study document souping, a strategy where documents are encoded independently, and their representations are pooled, via simple operations like averaging, into a single context state. This approach enables modular encoding and reuse without reprocessing the full input for each query. We demonstrate that finetuned Mamba2 models with souped representations achieve competitive or superior performance across multi-hop QA, sparse retrieval, and long-document reasoning tasks compared to the standard monolithic encoding approach. For example, on the RACE and QuALITY benchmarks for long document question answering, this method substantially outperforms a traditional concatenation approach. Crucially, this modular design scales to hundreds of documents while delivering substantial savings in inference cost, unlocking new possibilities for large-scale corpus reasoning.
title Studying the Soupability of Documents in State Space Models
topic Computation and Language
Computational Engineering, Finance, and Science
Machine Learning
url https://arxiv.org/abs/2505.24033