Leveraging Large Vision Model for Multi-UAV Co-perception in Low-Altitude Wireless Networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Yunting, Wang, Jiacheng, Zhang, Ruichen, Zhao, Changyuan, Liu, Yinqiu, Niyato, Dusit, Yu, Liang, Zhou, Haibo, Kim, Dong In
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911524415602688
author Xu, Yunting
Wang, Jiacheng
Zhang, Ruichen
Zhao, Changyuan
Liu, Yinqiu
Niyato, Dusit
Yu, Liang
Zhou, Haibo
Kim, Dong In
author_facet Xu, Yunting
Wang, Jiacheng
Zhang, Ruichen
Zhao, Changyuan
Liu, Yinqiu
Niyato, Dusit
Yu, Liang
Zhou, Haibo
Kim, Dong In
contents Multi-uncrewed aerial vehicle (UAV) cooperative perception has emerged as a promising paradigm for diverse low-altitude economy applications, where complementary multi-view observations are leveraged to enhance perception performance via wireless communications. However, the massive visual data generated by multiple UAVs poses significant challenges in terms of communication latency and resource efficiency. To address these challenges, this paper proposes a communication-efficient cooperative perception framework, termed Base-Station-Helped UAV (BHU), which reduces communication overhead while enhancing perception performance. Specifically, we employ a Top-K selection mechanism to identify the most informative pixels from UAV-captured RGB images, enabling sparsified visual transmission with reduced data volume and latency. The sparsified images are transmitted to a ground server via multi-user MIMO (MU-MIMO), where a Swin-large-based MaskDINO encoder extracts bird's-eye-view (BEV) features and performs cooperative feature fusion for ground vehicle perception. Furthermore, we develop a diffusion model-based deep reinforcement learning (DRL) algorithm to jointly select cooperative UAVs, sparsification ratios, and precoding matrices, achieving a balance between communication efficiency and perception utility. Simulation results on the Air-Co-Pred dataset demonstrate that, compared with traditional CNN-based BEV fusion baselines, the proposed BHU framework improves perception performance by over 5% while reducing communication overhead by 85%, providing an effective solution for multi-UAV cooperative perception under resource-constrained wireless environments.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16927
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Leveraging Large Vision Model for Multi-UAV Co-perception in Low-Altitude Wireless Networks
Xu, Yunting
Wang, Jiacheng
Zhang, Ruichen
Zhao, Changyuan
Liu, Yinqiu
Niyato, Dusit
Yu, Liang
Zhou, Haibo
Kim, Dong In
Computer Vision and Pattern Recognition
Image and Video Processing
Multi-uncrewed aerial vehicle (UAV) cooperative perception has emerged as a promising paradigm for diverse low-altitude economy applications, where complementary multi-view observations are leveraged to enhance perception performance via wireless communications. However, the massive visual data generated by multiple UAVs poses significant challenges in terms of communication latency and resource efficiency. To address these challenges, this paper proposes a communication-efficient cooperative perception framework, termed Base-Station-Helped UAV (BHU), which reduces communication overhead while enhancing perception performance. Specifically, we employ a Top-K selection mechanism to identify the most informative pixels from UAV-captured RGB images, enabling sparsified visual transmission with reduced data volume and latency. The sparsified images are transmitted to a ground server via multi-user MIMO (MU-MIMO), where a Swin-large-based MaskDINO encoder extracts bird's-eye-view (BEV) features and performs cooperative feature fusion for ground vehicle perception. Furthermore, we develop a diffusion model-based deep reinforcement learning (DRL) algorithm to jointly select cooperative UAVs, sparsification ratios, and precoding matrices, achieving a balance between communication efficiency and perception utility. Simulation results on the Air-Co-Pred dataset demonstrate that, compared with traditional CNN-based BEV fusion baselines, the proposed BHU framework improves perception performance by over 5% while reducing communication overhead by 85%, providing an effective solution for multi-UAV cooperative perception under resource-constrained wireless environments.
title Leveraging Large Vision Model for Multi-UAV Co-perception in Low-Altitude Wireless Networks
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2603.16927