Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Guo, Lu, Lidong, Liu, Yicheng, Dong, Liangrui, Zou, Lidong, Lv, Jixin, Li, Zhenquan, Mao, Xinyi, Pei, Baoqi, Wang, Shihao, Li, Zhiqi, Sapra, Karan, Liu, Fuxiao, Zheng, Yin-Dong, Huang, Yifei, Wang, Limin, Yu, Zhiding, Tao, Andrew, Liu, Guilin, Lu, Tong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915838110466048
author Chen, Guo
Lu, Lidong
Liu, Yicheng
Dong, Liangrui
Zou, Lidong
Lv, Jixin
Li, Zhenquan
Mao, Xinyi
Pei, Baoqi
Wang, Shihao
Li, Zhiqi
Sapra, Karan
Liu, Fuxiao
Zheng, Yin-Dong
Huang, Yifei
Wang, Limin
Yu, Zhiding
Tao, Andrew
Liu, Guilin
Lu, Tong
author_facet Chen, Guo
Lu, Lidong
Liu, Yicheng
Dong, Liangrui
Zou, Lidong
Lv, Jixin
Li, Zhenquan
Mao, Xinyi
Pei, Baoqi
Wang, Shihao
Li, Zhiqi
Sapra, Karan
Liu, Fuxiao
Zheng, Yin-Dong
Huang, Yifei
Wang, Limin
Yu, Zhiding
Tao, Andrew
Liu, Guilin
Lu, Tong
contents While datasets for video understanding have scaled to hour-long durations, they typically consist of densely concatenated clips that differ from natural, unscripted daily life. To bridge this gap, we introduce MM-Lifelong, a dataset designed for Multimodal Lifelong Understanding. Comprising 181.1 hours of footage, it is structured across Day, Week, and Month scales to capture varying temporal densities. Extensive evaluations reveal two critical failure modes in current paradigms: end-to-end MLLMs suffer from a Working Memory Bottleneck due to context saturation, while representative agentic baselines experience Global Localization Collapse when navigating sparse, month-long timelines. To address this, we propose the Recursive Multimodal Agent (ReMA), which employs dynamic memory management to iteratively update a recursive belief state, significantly outperforming existing methods. Finally, we establish dataset splits designed to isolate temporal and domain biases, providing a rigorous foundation for future research in supervised learning and out-of-distribution generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05484
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
Chen, Guo
Lu, Lidong
Liu, Yicheng
Dong, Liangrui
Zou, Lidong
Lv, Jixin
Li, Zhenquan
Mao, Xinyi
Pei, Baoqi
Wang, Shihao
Li, Zhiqi
Sapra, Karan
Liu, Fuxiao
Zheng, Yin-Dong
Huang, Yifei
Wang, Limin
Yu, Zhiding
Tao, Andrew
Liu, Guilin
Lu, Tong
Computer Vision and Pattern Recognition
While datasets for video understanding have scaled to hour-long durations, they typically consist of densely concatenated clips that differ from natural, unscripted daily life. To bridge this gap, we introduce MM-Lifelong, a dataset designed for Multimodal Lifelong Understanding. Comprising 181.1 hours of footage, it is structured across Day, Week, and Month scales to capture varying temporal densities. Extensive evaluations reveal two critical failure modes in current paradigms: end-to-end MLLMs suffer from a Working Memory Bottleneck due to context saturation, while representative agentic baselines experience Global Localization Collapse when navigating sparse, month-long timelines. To address this, we propose the Recursive Multimodal Agent (ReMA), which employs dynamic memory management to iteratively update a recursive belief state, significantly outperforming existing methods. Finally, we establish dataset splits designed to isolate temporal and domain biases, providing a rigorous foundation for future research in supervised learning and out-of-distribution generalization.
title Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.05484