Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Jinzheng, Xu, Yong, Qian, Xinyuan, Berghi, Davide, Wu, Peipei, Cui, Meng, Sun, Jianyuan, Jackson, Philip J. B., Wang, Wenwu
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908315551793152
author Zhao, Jinzheng
Xu, Yong
Qian, Xinyuan
Berghi, Davide
Wu, Peipei
Cui, Meng
Sun, Jianyuan
Jackson, Philip J. B.
Wang, Wenwu
author_facet Zhao, Jinzheng
Xu, Yong
Qian, Xinyuan
Berghi, Davide
Wu, Peipei
Cui, Meng
Sun, Jianyuan
Jackson, Philip J. B.
Wang, Wenwu
contents Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide complementary information for localization and tracking. With audio and visual information, the Bayesian-based filter and deep learning-based methods can solve the problem of data association, audio-visual fusion and track management. In this paper, we conduct a comprehensive overview of audio-visual speaker tracking. To our knowledge, this is the first extensive survey over the past five years. We introduce the family of Bayesian filters and summarize the methods for obtaining audio-visual measurements. In addition, the existing trackers and their performance on the AV16.3 dataset are summarized. In the past few years, deep learning techniques have thrived, which also boost the development of audio-visual speaker tracking. The influence of deep learning techniques in terms of measurement extraction and state estimation is also discussed. Finally, we discuss the connections between audio-visual speaker tracking and other areas such as speech separation and distributed speaker tracking.
format Preprint
id arxiv_https___arxiv_org_abs_2310_14778
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
Zhao, Jinzheng
Xu, Yong
Qian, Xinyuan
Berghi, Davide
Wu, Peipei
Cui, Meng
Sun, Jianyuan
Jackson, Philip J. B.
Wang, Wenwu
Multimedia
Sound
Audio and Speech Processing
Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide complementary information for localization and tracking. With audio and visual information, the Bayesian-based filter and deep learning-based methods can solve the problem of data association, audio-visual fusion and track management. In this paper, we conduct a comprehensive overview of audio-visual speaker tracking. To our knowledge, this is the first extensive survey over the past five years. We introduce the family of Bayesian filters and summarize the methods for obtaining audio-visual measurements. In addition, the existing trackers and their performance on the AV16.3 dataset are summarized. In the past few years, deep learning techniques have thrived, which also boost the development of audio-visual speaker tracking. The influence of deep learning techniques in terms of measurement extraction and state estimation is also discussed. Finally, we discuss the connections between audio-visual speaker tracking and other areas such as speech separation and distributed speaker tracking.
title Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
topic Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2310.14778