Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haque, Kazi Nazmul, Rana, Rajib, Jarin, Tasnim, Schuller Jr, Bjorn W.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909410318614528
author Haque, Kazi Nazmul
Rana, Rajib
Jarin, Tasnim
Schuller Jr, Bjorn W.
author_facet Haque, Kazi Nazmul
Rana, Rajib
Jarin, Tasnim
Schuller Jr, Bjorn W.
contents This study explores the field of audio classification from raw waveform using Convolutional Neural Networks (CNNs), a method that eliminates the need for extracting specialised features in the pre-processing step. Unlike recent trends in literature, which often focuses on designing frontends or filters for only the initial layers of CNNs, our research introduces the Cosine Convolutional Neural Network (CosCovNN) replacing the traditional CNN filters with Cosine filters. The CosCovNN surpasses the accuracy of the equivalent CNN architectures with approximately $77\%$ less parameters. Our research further progresses with the development of an augmented CosCovNN named Vector Quantised Cosine Convolutional Neural Network with Memory (VQCCM), incorporating a memory and vector quantisation layer VQCCM achieves state-of-the-art (SOTA) performance across five different datasets in comparison with existing literature. Our findings show that cosine filters can greatly improve the efficiency and accuracy of CNNs in raw audio classification.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00312
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)
Haque, Kazi Nazmul
Rana, Rajib
Jarin, Tasnim
Schuller Jr, Bjorn W.
Sound
Artificial Intelligence
Audio and Speech Processing
This study explores the field of audio classification from raw waveform using Convolutional Neural Networks (CNNs), a method that eliminates the need for extracting specialised features in the pre-processing step. Unlike recent trends in literature, which often focuses on designing frontends or filters for only the initial layers of CNNs, our research introduces the Cosine Convolutional Neural Network (CosCovNN) replacing the traditional CNN filters with Cosine filters. The CosCovNN surpasses the accuracy of the equivalent CNN architectures with approximately $77\%$ less parameters. Our research further progresses with the development of an augmented CosCovNN named Vector Quantised Cosine Convolutional Neural Network with Memory (VQCCM), incorporating a memory and vector quantisation layer VQCCM achieves state-of-the-art (SOTA) performance across five different datasets in comparison with existing literature. Our findings show that cosine filters can greatly improve the efficiency and accuracy of CNNs in raw audio classification.
title Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2412.00312