Dual Latent Memory for Visual Multi-agent System

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Xinlei, Xu, Chengming, Chen, Zhangquan, Yin, Bo, Yang, Cheng, He, Yongbo, Hu, Yihao, Zhang, Jiangning, Tan, Cheng, Hu, Xiaobin, Yan, Shuicheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908801676869632
author Yu, Xinlei
Xu, Chengming
Chen, Zhangquan
Yin, Bo
Yang, Cheng
He, Yongbo
Hu, Yihao
Zhang, Jiangning
Tan, Cheng
Hu, Xiaobin
Yan, Shuicheng
author_facet Yu, Xinlei
Xu, Chengming
Chen, Zhangquan
Yin, Bo
Yang, Cheng
He, Yongbo
Hu, Yihao
Zhang, Jiangning
Tan, Cheng
Hu, Xiaobin
Yan, Shuicheng
contents While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs. We attribute this failure to the information bottleneck inherent in text-centric communication, where converting perceptual and thinking trajectories into discrete natural language inevitably induces semantic loss. To this end, we propose L$^{2}$-VMAS, a novel model-agnostic framework that enables inter-agent collaboration with dual latent memories. Furthermore, we decouple the perception and thinking while dynamically synthesizing dual latent memories. Additionally, we introduce an entropy-driven proactive triggering that replaces passive information transmission with efficient, on-demand memory access. Extensive experiments among backbones, sizes, and multi-agent structures demonstrate that our method effectively breaks the "scaling wall" with superb scalability, improving average accuracy by 2.7-5.4% while reducing token usage by 21.3-44.8%. Codes: https://github.com/YU-deep/L2-VMAS.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00471
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dual Latent Memory for Visual Multi-agent System
Yu, Xinlei
Xu, Chengming
Chen, Zhangquan
Yin, Bo
Yang, Cheng
He, Yongbo
Hu, Yihao
Zhang, Jiangning
Tan, Cheng
Hu, Xiaobin
Yan, Shuicheng
Artificial Intelligence
Computer Vision and Pattern Recognition
While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs. We attribute this failure to the information bottleneck inherent in text-centric communication, where converting perceptual and thinking trajectories into discrete natural language inevitably induces semantic loss. To this end, we propose L$^{2}$-VMAS, a novel model-agnostic framework that enables inter-agent collaboration with dual latent memories. Furthermore, we decouple the perception and thinking while dynamically synthesizing dual latent memories. Additionally, we introduce an entropy-driven proactive triggering that replaces passive information transmission with efficient, on-demand memory access. Extensive experiments among backbones, sizes, and multi-agent structures demonstrate that our method effectively breaks the "scaling wall" with superb scalability, improving average accuracy by 2.7-5.4% while reducing token usage by 21.3-44.8%. Codes: https://github.com/YU-deep/L2-VMAS.
title Dual Latent Memory for Visual Multi-agent System
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.00471