Cross-Modal State-Space Graph Reasoning for Structured Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Hannah, Martinez, Sofia, Lee, Jason
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911083072061440
author Kim, Hannah
Martinez, Sofia
Lee, Jason
author_facet Kim, Hannah
Martinez, Sofia
Lee, Jason
contents The ability to extract compact, meaningful summaries from large-scale and multimodal data is critical for numerous applications, ranging from video analytics to medical reports. Prior methods in cross-modal summarization have often suffered from high computational overheads and limited interpretability. In this paper, we propose a \textit{Cross-Modal State-Space Graph Reasoning} (\textbf{CSS-GR}) framework that incorporates a state-space model with graph-based message passing, inspired by prior work on efficient state-space models. Unlike existing approaches relying on purely sequential models, our method constructs a graph that captures inter- and intra-modal relationships, allowing more holistic reasoning over both textual and visual streams. We demonstrate that our approach significantly improves summarization quality and interpretability while maintaining computational efficiency, as validated on standard multimodal summarization benchmarks. We also provide a thorough ablation study to highlight the contributions of each component.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20988
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cross-Modal State-Space Graph Reasoning for Structured Summarization
Kim, Hannah
Martinez, Sofia
Lee, Jason
Computation and Language
Graphics
The ability to extract compact, meaningful summaries from large-scale and multimodal data is critical for numerous applications, ranging from video analytics to medical reports. Prior methods in cross-modal summarization have often suffered from high computational overheads and limited interpretability. In this paper, we propose a \textit{Cross-Modal State-Space Graph Reasoning} (\textbf{CSS-GR}) framework that incorporates a state-space model with graph-based message passing, inspired by prior work on efficient state-space models. Unlike existing approaches relying on purely sequential models, our method constructs a graph that captures inter- and intra-modal relationships, allowing more holistic reasoning over both textual and visual streams. We demonstrate that our approach significantly improves summarization quality and interpretability while maintaining computational efficiency, as validated on standard multimodal summarization benchmarks. We also provide a thorough ablation study to highlight the contributions of each component.
title Cross-Modal State-Space Graph Reasoning for Structured Summarization
topic Computation and Language
Graphics
url https://arxiv.org/abs/2503.20988