Advancing Audio Fingerprinting Accuracy Addressing Background Noise and Distortion Challenges

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kamuni, Navin, Chintala, Sathishkumar, Kunchakuri, Naveen, Narasimharaju, Jyothi Swaroop Arlagadda, Kumar, Venkat
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913372661874688
author Kamuni, Navin
Chintala, Sathishkumar
Kunchakuri, Naveen
Narasimharaju, Jyothi Swaroop Arlagadda
Kumar, Venkat
author_facet Kamuni, Navin
Chintala, Sathishkumar
Kunchakuri, Naveen
Narasimharaju, Jyothi Swaroop Arlagadda
Kumar, Venkat
contents Audio fingerprinting, exemplified by pioneers like Shazam, has transformed digital audio recognition. However, existing systems struggle with accuracy in challenging conditions, limiting broad applicability. This research proposes an AI and ML integrated audio fingerprinting algorithm to enhance accuracy. Built on the Dejavu Project's foundations, the study emphasizes real-world scenario simulations with diverse background noises and distortions. Signal processing, central to Dejavu's model, includes the Fast Fourier Transform, spectrograms, and peak extraction. The "constellation" concept and fingerprint hashing enable unique song identification. Performance evaluation attests to 100% accuracy within a 5-second audio input, with a system showcasing predictable matching speed for efficiency. Storage analysis highlights the critical space-speed trade-off for practical implementation. This research advances audio fingerprinting's adaptability, addressing challenges in varied environments and applications.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13957
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Advancing Audio Fingerprinting Accuracy Addressing Background Noise and Distortion Challenges
Kamuni, Navin
Chintala, Sathishkumar
Kunchakuri, Naveen
Narasimharaju, Jyothi Swaroop Arlagadda
Kumar, Venkat
Sound
Machine Learning
Audio and Speech Processing
Audio fingerprinting, exemplified by pioneers like Shazam, has transformed digital audio recognition. However, existing systems struggle with accuracy in challenging conditions, limiting broad applicability. This research proposes an AI and ML integrated audio fingerprinting algorithm to enhance accuracy. Built on the Dejavu Project's foundations, the study emphasizes real-world scenario simulations with diverse background noises and distortions. Signal processing, central to Dejavu's model, includes the Fast Fourier Transform, spectrograms, and peak extraction. The "constellation" concept and fingerprint hashing enable unique song identification. Performance evaluation attests to 100% accuracy within a 5-second audio input, with a system showcasing predictable matching speed for efficiency. Storage analysis highlights the critical space-speed trade-off for practical implementation. This research advances audio fingerprinting's adaptability, addressing challenges in varied environments and applications.
title Advancing Audio Fingerprinting Accuracy Addressing Background Noise and Distortion Challenges
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2402.13957