Dual Latent Memory for Visual Multi-agent System
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908801676869632 |
|---|---|
| author | Yu, Xinlei Xu, Chengming Chen, Zhangquan Yin, Bo Yang, Cheng He, Yongbo Hu, Yihao Zhang, Jiangning Tan, Cheng Hu, Xiaobin Yan, Shuicheng |
| author_facet | Yu, Xinlei Xu, Chengming Chen, Zhangquan Yin, Bo Yang, Cheng He, Yongbo Hu, Yihao Zhang, Jiangning Tan, Cheng Hu, Xiaobin Yan, Shuicheng |
| contents | While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs. We attribute this failure to the information bottleneck inherent in text-centric communication, where converting perceptual and thinking trajectories into discrete natural language inevitably induces semantic loss. To this end, we propose L$^{2}$-VMAS, a novel model-agnostic framework that enables inter-agent collaboration with dual latent memories. Furthermore, we decouple the perception and thinking while dynamically synthesizing dual latent memories. Additionally, we introduce an entropy-driven proactive triggering that replaces passive information transmission with efficient, on-demand memory access. Extensive experiments among backbones, sizes, and multi-agent structures demonstrate that our method effectively breaks the "scaling wall" with superb scalability, improving average accuracy by 2.7-5.4% while reducing token usage by 21.3-44.8%. Codes: https://github.com/YU-deep/L2-VMAS. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_00471 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Dual Latent Memory for Visual Multi-agent System Yu, Xinlei Xu, Chengming Chen, Zhangquan Yin, Bo Yang, Cheng He, Yongbo Hu, Yihao Zhang, Jiangning Tan, Cheng Hu, Xiaobin Yan, Shuicheng Artificial Intelligence Computer Vision and Pattern Recognition While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs. We attribute this failure to the information bottleneck inherent in text-centric communication, where converting perceptual and thinking trajectories into discrete natural language inevitably induces semantic loss. To this end, we propose L$^{2}$-VMAS, a novel model-agnostic framework that enables inter-agent collaboration with dual latent memories. Furthermore, we decouple the perception and thinking while dynamically synthesizing dual latent memories. Additionally, we introduce an entropy-driven proactive triggering that replaces passive information transmission with efficient, on-demand memory access. Extensive experiments among backbones, sizes, and multi-agent structures demonstrate that our method effectively breaks the "scaling wall" with superb scalability, improving average accuracy by 2.7-5.4% while reducing token usage by 21.3-44.8%. Codes: https://github.com/YU-deep/L2-VMAS. |
| title | Dual Latent Memory for Visual Multi-agent System |
| topic | Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.00471 |