What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Su, Wanfang, Chen, Lixing, Bai, Yang, Lin, Xi, Li, Gaolei, Qu, Zhe, Zhou, Pan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913266173739008
author Su, Wanfang
Chen, Lixing
Bai, Yang
Lin, Xi
Li, Gaolei
Qu, Zhe
Zhou, Pan
author_facet Su, Wanfang
Chen, Lixing
Bai, Yang
Lin, Xi
Li, Gaolei
Qu, Zhe
Zhou, Pan
contents Multi-agent perception (MAP) allows autonomous systems to understand complex environments by interpreting data from multiple sources. This paper investigates intermediate collaboration for MAP with a specific focus on exploring "good" properties of collaborative view (i.e., post-collaboration feature) and its underlying relationship to individual views (i.e., pre-collaboration features), which were treated as an opaque procedure by most existing works. We propose a novel framework named CMiMC (Contrastive Mutual Information Maximization for Collaborative Perception) for intermediate collaboration. The core philosophy of CMiMC is to preserve discriminative information of individual views in the collaborative view by maximizing mutual information between pre- and post-collaboration features while enhancing the efficacy of collaborative views by minimizing the loss function of downstream tasks. In particular, we define multi-view mutual information (MVMI) for intermediate collaboration that evaluates correlations between collaborative views and individual views on both global and local scales. We establish CMiMNet based on multi-view contrastive learning to realize estimation and maximization of MVMI, which assists the training of a collaboration encoder for voxel-level feature fusion. We evaluate CMiMC on V2X-Sim 1.0, and it improves the SOTA average precision by 3.08% and 4.44% at 0.5 and 0.7 IoU (Intersection-over-Union) thresholds, respectively. In addition, CMiMC can reduce communication volume to 1/32 while achieving performance comparable to SOTA. Code and Appendix are released at https://github.com/77SWF/CMiMC.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10068
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception
Su, Wanfang
Chen, Lixing
Bai, Yang
Lin, Xi
Li, Gaolei
Qu, Zhe
Zhou, Pan
Computer Vision and Pattern Recognition
Multiagent Systems
Multi-agent perception (MAP) allows autonomous systems to understand complex environments by interpreting data from multiple sources. This paper investigates intermediate collaboration for MAP with a specific focus on exploring "good" properties of collaborative view (i.e., post-collaboration feature) and its underlying relationship to individual views (i.e., pre-collaboration features), which were treated as an opaque procedure by most existing works. We propose a novel framework named CMiMC (Contrastive Mutual Information Maximization for Collaborative Perception) for intermediate collaboration. The core philosophy of CMiMC is to preserve discriminative information of individual views in the collaborative view by maximizing mutual information between pre- and post-collaboration features while enhancing the efficacy of collaborative views by minimizing the loss function of downstream tasks. In particular, we define multi-view mutual information (MVMI) for intermediate collaboration that evaluates correlations between collaborative views and individual views on both global and local scales. We establish CMiMNet based on multi-view contrastive learning to realize estimation and maximization of MVMI, which assists the training of a collaboration encoder for voxel-level feature fusion. We evaluate CMiMC on V2X-Sim 1.0, and it improves the SOTA average precision by 3.08% and 4.44% at 0.5 and 0.7 IoU (Intersection-over-Union) thresholds, respectively. In addition, CMiMC can reduce communication volume to 1/32 while achieving performance comparable to SOTA. Code and Appendix are released at https://github.com/77SWF/CMiMC.
title What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception
topic Computer Vision and Pattern Recognition
Multiagent Systems
url https://arxiv.org/abs/2403.10068