AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866929644515622912 |
|---|---|
| author | Xiao, Zhenyuan Yang, Yizhuo Xu, Guili Zeng, Xianglong Yuan, Shenghai |
| author_facet | Xiao, Zhenyuan Yang, Yizhuo Xu, Guili Zeng, Xianglong Yuan, Shenghai |
| contents | The increasing use of compact UAVs has created significant threats to public safety, while traditional drone detection systems are often bulky and costly. To address these challenges, we propose AV-DTEC, a lightweight self-supervised audio-visual fusion-based anti-UAV system. AV-DTEC is trained using self-supervised learning with labels generated by LiDAR, and it simultaneously learns audio and visual features through a parallel selective state-space model. With the learned features, a specially designed plug-and-play primary-auxiliary feature enhancement module integrates visual features into audio features for better robustness in cross-lighting conditions. To reduce reliance on auxiliary features and align modalities, we propose a teacher-student model that adaptively adjusts the weighting of visual features. AV-DTEC demonstrates exceptional accuracy and effectiveness in real-world multi-modality data. The code and trained models are publicly accessible on GitHub
\url{https://github.com/AmazingDay1/AV-DETC}. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_16928 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification Xiao, Zhenyuan Yang, Yizhuo Xu, Guili Zeng, Xianglong Yuan, Shenghai Sound Computer Vision and Pattern Recognition Multimedia Audio and Speech Processing The increasing use of compact UAVs has created significant threats to public safety, while traditional drone detection systems are often bulky and costly. To address these challenges, we propose AV-DTEC, a lightweight self-supervised audio-visual fusion-based anti-UAV system. AV-DTEC is trained using self-supervised learning with labels generated by LiDAR, and it simultaneously learns audio and visual features through a parallel selective state-space model. With the learned features, a specially designed plug-and-play primary-auxiliary feature enhancement module integrates visual features into audio features for better robustness in cross-lighting conditions. To reduce reliance on auxiliary features and align modalities, we propose a teacher-student model that adaptively adjusts the weighting of visual features. AV-DTEC demonstrates exceptional accuracy and effectiveness in real-world multi-modality data. The code and trained models are publicly accessible on GitHub \url{https://github.com/AmazingDay1/AV-DETC}. |
| title | AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification |
| topic | Sound Computer Vision and Pattern Recognition Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2412.16928 |