TACNET: Temporal Audio Source Counting Network
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866916534480273408 |
|---|---|
| author | Ahmadnejad, Amirreza Darviishani, Ahmad Mahmmodian Asadi, Mohmmad Mehrdad Saffariyeh, Sajjad Yousef, Pedram Fatemizadeh, Emad |
| author_facet | Ahmadnejad, Amirreza Darviishani, Ahmad Mahmmodian Asadi, Mohmmad Mehrdad Saffariyeh, Sajjad Yousef, Pedram Fatemizadeh, Emad |
| contents | In this paper, we introduce the Temporal Audio Source Counting Network (TaCNet), an innovative architecture that addresses limitations in audio source counting tasks. TaCNet operates directly on raw audio inputs, eliminating complex preprocessing steps and simplifying the workflow. Notably, it excels in real-time speaker counting, even with truncated input windows. Our extensive evaluation, conducted using the LibriCount dataset, underscores TaCNet's exceptional performance, positioning it as a state-of-the-art solution for audio source counting tasks. With an average accuracy of 74.18 percentage over 11 classes, TaCNet demonstrates its effectiveness across diverse scenarios, including applications involving Chinese and Persian languages. This cross-lingual adaptability highlights its versatility and potential impact. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_02369 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | TACNET: Temporal Audio Source Counting Network Ahmadnejad, Amirreza Darviishani, Ahmad Mahmmodian Asadi, Mohmmad Mehrdad Saffariyeh, Sajjad Yousef, Pedram Fatemizadeh, Emad Sound Artificial Intelligence Machine Learning Audio and Speech Processing In this paper, we introduce the Temporal Audio Source Counting Network (TaCNet), an innovative architecture that addresses limitations in audio source counting tasks. TaCNet operates directly on raw audio inputs, eliminating complex preprocessing steps and simplifying the workflow. Notably, it excels in real-time speaker counting, even with truncated input windows. Our extensive evaluation, conducted using the LibriCount dataset, underscores TaCNet's exceptional performance, positioning it as a state-of-the-art solution for audio source counting tasks. With an average accuracy of 74.18 percentage over 11 classes, TaCNet demonstrates its effectiveness across diverse scenarios, including applications involving Chinese and Persian languages. This cross-lingual adaptability highlights its versatility and potential impact. |
| title | TACNET: Temporal Audio Source Counting Network |
| topic | Sound Artificial Intelligence Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2311.02369 |