Advanced Framework for Animal Sound Classification With Features Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Qiang, Chen, Xiuying, Ma, Changsheng, Duarte, Carlos M., Zhang, Xiangliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911943998046208
author Yang, Qiang
Chen, Xiuying
Ma, Changsheng
Duarte, Carlos M.
Zhang, Xiangliang
author_facet Yang, Qiang
Chen, Xiuying
Ma, Changsheng
Duarte, Carlos M.
Zhang, Xiangliang
contents The automatic classification of animal sounds presents an enduring challenge in bioacoustics, owing to the diverse statistical properties of sound signals, variations in recording equipment, and prevalent low Signal-to-Noise Ratio (SNR) conditions. Deep learning models like Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) have excelled in human speech recognition but have not been effectively tailored to the intricate nature of animal sounds, which exhibit substantial diversity even within the same domain. We propose an automated classification framework applicable to general animal sound classification. Our approach first optimizes audio features from Mel-frequency cepstral coefficients (MFCC) including feature rearrangement and feature reduction. It then uses the optimized features for the deep learning model, i.e., an attention-based Bidirectional LSTM (Bi-LSTM), to extract deep semantic features for sound classification. We also contribute an animal sound benchmark dataset encompassing oceanic animals and birds1. Extensive experimentation with real-world datasets demonstrates that our approach consistently outperforms baseline methods by over 25% in precision, recall, and accuracy, promising advancements in animal sound classification.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03440
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Advanced Framework for Animal Sound Classification With Features Optimization
Yang, Qiang
Chen, Xiuying
Ma, Changsheng
Duarte, Carlos M.
Zhang, Xiangliang
Sound
Machine Learning
Audio and Speech Processing
The automatic classification of animal sounds presents an enduring challenge in bioacoustics, owing to the diverse statistical properties of sound signals, variations in recording equipment, and prevalent low Signal-to-Noise Ratio (SNR) conditions. Deep learning models like Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) have excelled in human speech recognition but have not been effectively tailored to the intricate nature of animal sounds, which exhibit substantial diversity even within the same domain. We propose an automated classification framework applicable to general animal sound classification. Our approach first optimizes audio features from Mel-frequency cepstral coefficients (MFCC) including feature rearrangement and feature reduction. It then uses the optimized features for the deep learning model, i.e., an attention-based Bidirectional LSTM (Bi-LSTM), to extract deep semantic features for sound classification. We also contribute an animal sound benchmark dataset encompassing oceanic animals and birds1. Extensive experimentation with real-world datasets demonstrates that our approach consistently outperforms baseline methods by over 25% in precision, recall, and accuracy, promising advancements in animal sound classification.
title Advanced Framework for Animal Sound Classification With Features Optimization
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2407.03440