Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Kerui, Wang, Jinglu, Zhang, Jianrong, Li, Ming, Lu, Yan, Fan, Hehe
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909007620341760
author Chen, Kerui
Wang, Jinglu
Zhang, Jianrong
Li, Ming
Lu, Yan
Fan, Hehe
author_facet Chen, Kerui
Wang, Jinglu
Zhang, Jianrong
Li, Ming
Lu, Yan
Fan, Hehe
contents Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existing agentic methods mitigate this via rule-based preprocessing, yet often suffer from information loss, high cost, and reliance on textual intermediates. We propose MACF, an end-to-end Multi-Agent Collaboration Framework that decouples per-agent perception budgets from global video complexity, enabling scalable video understanding while preserving visual fidelity. MACF partitions videos into segments for locally budgeted agents and enables holistic reasoning via an agent-native latent communication protocol. Each agent encodes partial observations into compact, task-sufficient tokens in a shared embedding space, allowing efficient and information-preserving collaboration by a central coordinator. We introduce a curriculum training strategy that progressively enforces semantic alignment, evidence summarization, and cross-agent coordination. Extensive experiments on diverse video understanding benchmarks show that MACF consistently outperforms state-of-the-art MLLMs and multi-agent systems under identical budget constraints, demonstrating the effectiveness of our latent collaboration for scalable video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00444
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
Chen, Kerui
Wang, Jinglu
Zhang, Jianrong
Li, Ming
Lu, Yan
Fan, Hehe
Computer Vision and Pattern Recognition
Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existing agentic methods mitigate this via rule-based preprocessing, yet often suffer from information loss, high cost, and reliance on textual intermediates. We propose MACF, an end-to-end Multi-Agent Collaboration Framework that decouples per-agent perception budgets from global video complexity, enabling scalable video understanding while preserving visual fidelity. MACF partitions videos into segments for locally budgeted agents and enables holistic reasoning via an agent-native latent communication protocol. Each agent encodes partial observations into compact, task-sufficient tokens in a shared embedding space, allowing efficient and information-preserving collaboration by a central coordinator. We introduce a curriculum training strategy that progressively enforces semantic alignment, evidence summarization, and cross-agent coordination. Extensive experiments on diverse video understanding benchmarks show that MACF consistently outperforms state-of-the-art MLLMs and multi-agent systems under identical budget constraints, demonstrating the effectiveness of our latent collaboration for scalable video understanding.
title Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.00444