Unified AI for Accurate Audio Anomaly Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khaleghpour, Hamideh, McKinney, Brett
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909628174958592
author Khaleghpour, Hamideh
McKinney, Brett
author_facet Khaleghpour, Hamideh
McKinney, Brett
contents This paper presents a unified AI framework for high-accuracy audio anomaly detection by integrating advanced noise reduction, feature extraction, and machine learning modeling techniques. The approach combines spectral subtraction and adaptive filtering to enhance audio quality, followed by feature extraction using traditional methods like MFCCs and deep embeddings from pre-trained models such as OpenL3. The modeling pipeline incorporates classical models (SVM, Random Forest), deep learning architectures (CNNs), and ensemble methods to boost robustness and accuracy. Evaluated on benchmark datasets including TORGO and LibriSpeech, the proposed framework demonstrates superior performance in precision, recall, and classification of slurred vs. normal speech. This work addresses challenges in noisy environments and real-time applications and provides a scalable solution for audio-based anomaly detection.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unified AI for Accurate Audio Anomaly Detection
Khaleghpour, Hamideh
McKinney, Brett
Sound
Machine Learning
Audio and Speech Processing
This paper presents a unified AI framework for high-accuracy audio anomaly detection by integrating advanced noise reduction, feature extraction, and machine learning modeling techniques. The approach combines spectral subtraction and adaptive filtering to enhance audio quality, followed by feature extraction using traditional methods like MFCCs and deep embeddings from pre-trained models such as OpenL3. The modeling pipeline incorporates classical models (SVM, Random Forest), deep learning architectures (CNNs), and ensemble methods to boost robustness and accuracy. Evaluated on benchmark datasets including TORGO and LibriSpeech, the proposed framework demonstrates superior performance in precision, recall, and classification of slurred vs. normal speech. This work addresses challenges in noisy environments and real-time applications and provides a scalable solution for audio-based anomaly detection.
title Unified AI for Accurate Audio Anomaly Detection
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2505.23781