MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Minghao, Jiao, Qingyue, Shi, Zeru, Quan, Yihao, Zhang, Boxuan, Li, Danrui, Che, Liwei, Xu, Wujiang, Liu, Shilong, Liu, Zirui, Kapadia, Mubbasir, Pavlovic, Vladimir, Liu, Jiang, Wang, Mengdi, Shi, Yiyu, Metaxas, Dimitris N., Tang, Ruixiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910222397734912
author Guo, Minghao
Jiao, Qingyue
Shi, Zeru
Quan, Yihao
Zhang, Boxuan
Li, Danrui
Che, Liwei
Xu, Wujiang
Liu, Shilong
Liu, Zirui
Kapadia, Mubbasir
Pavlovic, Vladimir
Liu, Jiang
Wang, Mengdi
Shi, Yiyu
Metaxas, Dimitris N.
Tang, Ruixiang
author_facet Guo, Minghao
Jiao, Qingyue
Shi, Zeru
Quan, Yihao
Zhang, Boxuan
Li, Danrui
Che, Liwei
Xu, Wujiang
Liu, Shilong
Liu, Zirui
Kapadia, Mubbasir
Pavlovic, Vladimir
Liu, Jiang
Wang, Mengdi
Shi, Yiyu
Metaxas, Dimitris N.
Tang, Ruixiang
contents Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later reasoning. In prior work, many visually grounded questions can be answered using only captions or textual traces, allowing answers to be inferred without preserving the fine-grained visual evidence. Meanwhile, harder cases that require reasoning over changing visual states are largely absent. Therefore, we introduce MemEye, a framework that evaluates memory capabilities from two dimensions: one measures the granularity of decisive visual evidence (from scene-level to pixel-level evidence), and the other measures how retrieved evidence must be used (from single evidence to evolutionary synthesis). Under this framework, we construct a new benchmark across 8 life-scenario tasks, with ablation-driven validation gates for assessing answerability, shortcut resistance, visual necessity, and reasoning structure. By evaluating 13 memory methods across 4 VLM backbones, we show that current architectures still struggle to preserve fine-grained visual details and reason about state changes over time. Our findings show that long-term multimodal memory depends on evidence routing, temporal tracking, and detail extraction.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15128
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
Guo, Minghao
Jiao, Qingyue
Shi, Zeru
Quan, Yihao
Zhang, Boxuan
Li, Danrui
Che, Liwei
Xu, Wujiang
Liu, Shilong
Liu, Zirui
Kapadia, Mubbasir
Pavlovic, Vladimir
Liu, Jiang
Wang, Mengdi
Shi, Yiyu
Metaxas, Dimitris N.
Tang, Ruixiang
Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later reasoning. In prior work, many visually grounded questions can be answered using only captions or textual traces, allowing answers to be inferred without preserving the fine-grained visual evidence. Meanwhile, harder cases that require reasoning over changing visual states are largely absent. Therefore, we introduce MemEye, a framework that evaluates memory capabilities from two dimensions: one measures the granularity of decisive visual evidence (from scene-level to pixel-level evidence), and the other measures how retrieved evidence must be used (from single evidence to evolutionary synthesis). Under this framework, we construct a new benchmark across 8 life-scenario tasks, with ablation-driven validation gates for assessing answerability, shortcut resistance, visual necessity, and reasoning structure. By evaluating 13 memory methods across 4 VLM backbones, we show that current architectures still struggle to preserve fine-grained visual details and reason about state changes over time. Our findings show that long-term multimodal memory depends on evidence routing, temporal tracking, and detail extraction.
title MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
topic Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2605.15128