AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xiao, Zhenyuan, Yang, Yizhuo, Xu, Guili, Zeng, Xianglong, Yuan, Shenghai
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929644515622912
author Xiao, Zhenyuan
Yang, Yizhuo
Xu, Guili
Zeng, Xianglong
Yuan, Shenghai
author_facet Xiao, Zhenyuan
Yang, Yizhuo
Xu, Guili
Zeng, Xianglong
Yuan, Shenghai
contents The increasing use of compact UAVs has created significant threats to public safety, while traditional drone detection systems are often bulky and costly. To address these challenges, we propose AV-DTEC, a lightweight self-supervised audio-visual fusion-based anti-UAV system. AV-DTEC is trained using self-supervised learning with labels generated by LiDAR, and it simultaneously learns audio and visual features through a parallel selective state-space model. With the learned features, a specially designed plug-and-play primary-auxiliary feature enhancement module integrates visual features into audio features for better robustness in cross-lighting conditions. To reduce reliance on auxiliary features and align modalities, we propose a teacher-student model that adaptively adjusts the weighting of visual features. AV-DTEC demonstrates exceptional accuracy and effectiveness in real-world multi-modality data. The code and trained models are publicly accessible on GitHub \url{https://github.com/AmazingDay1/AV-DETC}.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16928
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification
Xiao, Zhenyuan
Yang, Yizhuo
Xu, Guili
Zeng, Xianglong
Yuan, Shenghai
Sound
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
The increasing use of compact UAVs has created significant threats to public safety, while traditional drone detection systems are often bulky and costly. To address these challenges, we propose AV-DTEC, a lightweight self-supervised audio-visual fusion-based anti-UAV system. AV-DTEC is trained using self-supervised learning with labels generated by LiDAR, and it simultaneously learns audio and visual features through a parallel selective state-space model. With the learned features, a specially designed plug-and-play primary-auxiliary feature enhancement module integrates visual features into audio features for better robustness in cross-lighting conditions. To reduce reliance on auxiliary features and align modalities, we propose a teacher-student model that adaptively adjusts the weighting of visual features. AV-DTEC demonstrates exceptional accuracy and effectiveness in real-world multi-modality data. The code and trained models are publicly accessible on GitHub \url{https://github.com/AmazingDay1/AV-DETC}.
title AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification
topic Sound
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2412.16928