Pretrained Conformers for Audio Fingerprinting and Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Altwlkany, Kemal, Selmanovic, Elmedin, Delalic, Sead
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908532699299840
author Altwlkany, Kemal
Selmanovic, Elmedin
Delalic, Sead
author_facet Altwlkany, Kemal
Selmanovic, Elmedin
Delalic, Sead
contents Conformers have shown great results in speech processing due to their ability to capture both local and global interactions. In this work, we utilize a self-supervised contrastive learning framework to train conformer-based encoders that are capable of generating unique embeddings for small segments of audio, generalizing well to previously unseen data. We achieve state-of-the-art results for audio retrieval tasks while using only 3 seconds of audio to generate embeddings. Our models are almost completely immune to temporal misalignments and achieve state-of-the-art results in cases of other audio distortions such as noise, reverb or extreme temporal stretching. Code and models are made publicly available and the results are easy to reproduce as we train and test using popular and freely available datasets of different sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11609
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pretrained Conformers for Audio Fingerprinting and Retrieval
Altwlkany, Kemal
Selmanovic, Elmedin
Delalic, Sead
Sound
Artificial Intelligence
Information Retrieval
Audio and Speech Processing
Conformers have shown great results in speech processing due to their ability to capture both local and global interactions. In this work, we utilize a self-supervised contrastive learning framework to train conformer-based encoders that are capable of generating unique embeddings for small segments of audio, generalizing well to previously unseen data. We achieve state-of-the-art results for audio retrieval tasks while using only 3 seconds of audio to generate embeddings. Our models are almost completely immune to temporal misalignments and achieve state-of-the-art results in cases of other audio distortions such as noise, reverb or extreme temporal stretching. Code and models are made publicly available and the results are easy to reproduce as we train and test using popular and freely available datasets of different sizes.
title Pretrained Conformers for Audio Fingerprinting and Retrieval
topic Sound
Artificial Intelligence
Information Retrieval
Audio and Speech Processing
url https://arxiv.org/abs/2508.11609