CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Rujia, Gao, Xiangbo, Xiang, Hao, Xu, Runsheng, Tu, Zhengzhong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918076603170816
author Wang, Rujia
Gao, Xiangbo
Xiang, Hao
Xu, Runsheng
Tu, Zhengzhong
author_facet Wang, Rujia
Gao, Xiangbo
Xiang, Hao
Xu, Runsheng
Tu, Zhengzhong
contents Multi-agent collaborative perception enhances each agent perceptual capabilities by sharing sensing information to cooperatively perform robot perception tasks. This approach has proven effective in addressing challenges such as sensor deficiencies, occlusions, and long-range perception. However, existing representative collaborative perception systems transmit intermediate feature maps, such as bird-eye view (BEV) representations, which contain a significant amount of non-critical information, leading to high communication bandwidth requirements. To enhance communication efficiency while preserving perception capability, we introduce CoCMT, an object-query-based collaboration framework that optimizes communication bandwidth by selectively extracting and transmitting essential features. Within CoCMT, we introduce the Efficient Query Transformer (EQFormer) to effectively fuse multi-agent object queries and implement a synergistic deep supervision to enhance the positive reinforcement between stages, leading to improved overall performance. Experiments on OPV2V and V2V4Real datasets show CoCMT outperforms state-of-the-art methods while drastically reducing communication needs. On V2V4Real, our model (Top-50 object queries) requires only 0.416 Mb bandwidth, 83 times less than SOTA methods, while improving AP70 by 1.1 percent. This efficiency breakthrough enables practical collaborative perception deployment in bandwidth-constrained environments without sacrificing detection accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13504
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
Wang, Rujia
Gao, Xiangbo
Xiang, Hao
Xu, Runsheng
Tu, Zhengzhong
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Robotics
Multi-agent collaborative perception enhances each agent perceptual capabilities by sharing sensing information to cooperatively perform robot perception tasks. This approach has proven effective in addressing challenges such as sensor deficiencies, occlusions, and long-range perception. However, existing representative collaborative perception systems transmit intermediate feature maps, such as bird-eye view (BEV) representations, which contain a significant amount of non-critical information, leading to high communication bandwidth requirements. To enhance communication efficiency while preserving perception capability, we introduce CoCMT, an object-query-based collaboration framework that optimizes communication bandwidth by selectively extracting and transmitting essential features. Within CoCMT, we introduce the Efficient Query Transformer (EQFormer) to effectively fuse multi-agent object queries and implement a synergistic deep supervision to enhance the positive reinforcement between stages, leading to improved overall performance. Experiments on OPV2V and V2V4Real datasets show CoCMT outperforms state-of-the-art methods while drastically reducing communication needs. On V2V4Real, our model (Top-50 object queries) requires only 0.416 Mb bandwidth, 83 times less than SOTA methods, while improving AP70 by 1.1 percent. This efficiency breakthrough enables practical collaborative perception deployment in bandwidth-constrained environments without sacrificing detection accuracy.
title CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2503.13504