SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chai, Huacan, Wang, Yukai, Yang, Yingxuan, Peng, Dan, Song, Yuanyi, Fu, Zhihui, Liu, Weiwen, Lin, Jianghao, Wang, Jun, Zhang, Weinan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910223575285760
author Chai, Huacan
Wang, Yukai
Yang, Yingxuan
Peng, Dan
Song, Yuanyi
Fu, Zhihui
Liu, Weiwen
Lin, Jianghao
Wang, Jun
Zhang, Weinan
author_facet Chai, Huacan
Wang, Yukai
Yang, Yingxuan
Peng, Dan
Song, Yuanyi
Fu, Zhihui
Liu, Weiwen
Lin, Jianghao
Wang, Jun
Zhang, Weinan
contents Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use evidence distributed across independently originated sources. We argue that source-distributed memory composition is an important and under-examined bottleneck in multimodal agent memory, especially when relevant evidence is fragmented across heterogeneous artifacts such as conversations, profiles, screenshots, tables, images, and documents. To address this gap, we introduce Source-distributed Multimodal Memory Benchmark(SMMBench), which measures whether agents can retrieve, align, and compose multimodal evidence scattered across multiple sources rather than reason within a single curated context. SMMBench evaluates four core capabilities: (1) cross-source multimodal reasoning; (2) conflict resolution; (3) preference reasoning; (4) memory-grounded action prediction. The benchmark contains 1877 samples grounded in 264 sources. Experiments on representative memory-style and retrieval-based baselines show that current systems still struggle on these capabilities, positioning source-distributed multimodal memory as an important and still under-evaluated challenge for multimodal agents. Our data are available at https://huggingface.co/datasets/HuacanChai/SMMBench.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15710
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory
Chai, Huacan
Wang, Yukai
Yang, Yingxuan
Peng, Dan
Song, Yuanyi
Fu, Zhihui
Liu, Weiwen
Lin, Jianghao
Wang, Jun
Zhang, Weinan
Computation and Language
Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use evidence distributed across independently originated sources. We argue that source-distributed memory composition is an important and under-examined bottleneck in multimodal agent memory, especially when relevant evidence is fragmented across heterogeneous artifacts such as conversations, profiles, screenshots, tables, images, and documents. To address this gap, we introduce Source-distributed Multimodal Memory Benchmark(SMMBench), which measures whether agents can retrieve, align, and compose multimodal evidence scattered across multiple sources rather than reason within a single curated context. SMMBench evaluates four core capabilities: (1) cross-source multimodal reasoning; (2) conflict resolution; (3) preference reasoning; (4) memory-grounded action prediction. The benchmark contains 1877 samples grounded in 264 sources. Experiments on representative memory-style and retrieval-based baselines show that current systems still struggle on these capabilities, positioning source-distributed multimodal memory as an important and still under-evaluated challenge for multimodal agents. Our data are available at https://huggingface.co/datasets/HuacanChai/SMMBench.
title SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory
topic Computation and Language
url https://arxiv.org/abs/2605.15710