Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization
Fuente:
arXiv
Guardado en:
| Autores principales: | Altwlkany, Kemal, Delalić, Sead, Alihodžić, Adis, Selmanović, Elmedin, Hasić, Damir |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pretrained Conformers for Audio Fingerprinting and Retrieval
por: Altwlkany, Kemal, et al.
Publicado: (2025)
por: Altwlkany, Kemal, et al.
Publicado: (2025)
A Recurrent Neural Network Approach to the Answering Machine Detection Problem
por: Altwlkany, Kemal, et al.
Publicado: (2024)
por: Altwlkany, Kemal, et al.
Publicado: (2024)
On the Language and Gender Biases in PSTN, VoIP and Neural Audio Codecs
por: Altwlkany, Kemal, et al.
Publicado: (2025)
por: Altwlkany, Kemal, et al.
Publicado: (2025)
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
PeakNetFP: Peak-based Neural Audio Fingerprinting Robust to Extreme Time Stretching
por: Cortès-Sebastià, Guillem, et al.
Publicado: (2025)
por: Cortès-Sebastià, Guillem, et al.
Publicado: (2025)
SECP: A Speech Enhancement-Based Curation Pipeline For Scalable Acquisition Of Clean Speech
por: Sabra, Adam, et al.
Publicado: (2024)
por: Sabra, Adam, et al.
Publicado: (2024)
Language-based Audio Retrieval with Co-Attention Networks
por: Sun, Haoran, et al.
Publicado: (2024)
por: Sun, Haoran, et al.
Publicado: (2024)
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
por: Xin, Yifei, et al.
Publicado: (2024)
por: Xin, Yifei, et al.
Publicado: (2024)
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval
por: Wang, Qian, et al.
Publicado: (2024)
por: Wang, Qian, et al.
Publicado: (2024)
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2024)
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2024)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
por: Lin, Chyi-Jiunn, et al.
Publicado: (2024)
por: Lin, Chyi-Jiunn, et al.
Publicado: (2024)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
por: Zhang, Chu Yuan, et al.
Publicado: (2023)
por: Zhang, Chu Yuan, et al.
Publicado: (2023)
Dissecting Temporal Understanding in Text-to-Audio Retrieval
por: Oncescu, Andreea-Maria, et al.
Publicado: (2024)
por: Oncescu, Andreea-Maria, et al.
Publicado: (2024)
Knowledge Distillation for Real-Time Classification of Early Media in Voice Communications
por: Altwlkany, Kemal, et al.
Publicado: (2024)
por: Altwlkany, Kemal, et al.
Publicado: (2024)
Contrastive and Transfer Learning for Effective Audio Fingerprinting through a Real-World Evaluation Protocol
por: Nikou, Christos, et al.
Publicado: (2025)
por: Nikou, Christos, et al.
Publicado: (2025)
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
por: O'Reilly, Patrick, et al.
Publicado: (2025)
por: O'Reilly, Patrick, et al.
Publicado: (2025)
LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
por: Salganik, Rebecca, et al.
Publicado: (2024)
por: Salganik, Rebecca, et al.
Publicado: (2024)
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
por: Kolehmainen, Jari, et al.
Publicado: (2024)
por: Kolehmainen, Jari, et al.
Publicado: (2024)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
por: Li, Jiaqi, et al.
Publicado: (2024)
por: Li, Jiaqi, et al.
Publicado: (2024)
GraFPrint: A GNN-Based Approach for Audio Identification
por: Bhattacharjee, Aditya, et al.
Publicado: (2024)
por: Bhattacharjee, Aditya, et al.
Publicado: (2024)
Enhancing Crowdsourced Audio for Text-to-Speech Models
por: Giraldo, José, et al.
Publicado: (2024)
por: Giraldo, José, et al.
Publicado: (2024)
Expressivity-aware Music Performance Retrieval using Mid-level Perceptual Features and Emotion Word Embeddings
por: Chowdhury, Shreyan, et al.
Publicado: (2024)
por: Chowdhury, Shreyan, et al.
Publicado: (2024)
Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance
por: Gong, Ziling, et al.
Publicado: (2026)
por: Gong, Ziling, et al.
Publicado: (2026)
Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
por: Deng, Yimin, et al.
Publicado: (2024)
por: Deng, Yimin, et al.
Publicado: (2024)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
por: Sun, Haoqin, et al.
Publicado: (2025)
por: Sun, Haoqin, et al.
Publicado: (2025)
Anchor-aware Deep Metric Learning for Audio-visual Retrieval
por: Zeng, Donghuo, et al.
Publicado: (2024)
por: Zeng, Donghuo, et al.
Publicado: (2024)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
por: Ogg, Mattson, et al.
Publicado: (2025)
por: Ogg, Mattson, et al.
Publicado: (2025)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
por: Lin, Zhaofeng, et al.
Publicado: (2024)
por: Lin, Zhaofeng, et al.
Publicado: (2024)
Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech
por: Daeglau, Mareike, et al.
Publicado: (2025)
por: Daeglau, Mareike, et al.
Publicado: (2025)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
por: Xue, Jinlong, et al.
Publicado: (2024)
por: Xue, Jinlong, et al.
Publicado: (2024)
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
por: Fu, Siyuan, et al.
Publicado: (2025)
por: Fu, Siyuan, et al.
Publicado: (2025)
Multi-Sample Dynamic Time Warping for Few-Shot Keyword Spotting
por: Wilkinghoff, Kevin, et al.
Publicado: (2024)
por: Wilkinghoff, Kevin, et al.
Publicado: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
por: Dang, Trung, et al.
Publicado: (2024)
por: Dang, Trung, et al.
Publicado: (2024)
Advances in Speech Separation: Techniques, Challenges, and Future Trends
por: Li, Kai, et al.
Publicado: (2025)
por: Li, Kai, et al.
Publicado: (2025)
SynHate: Detecting Hate Speech in Synthetic Deepfake Audio
por: Ranjan, Rishabh, et al.
Publicado: (2025)
por: Ranjan, Rishabh, et al.
Publicado: (2025)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
por: Zhang, Xueyao, et al.
Publicado: (2023)
por: Zhang, Xueyao, et al.
Publicado: (2023)
Audio Fingerprinting with Holographic Reduced Representations
por: Fujita, Yusuke, et al.
Publicado: (2024)
por: Fujita, Yusuke, et al.
Publicado: (2024)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
por: Wang, Shaowen, et al.
Publicado: (2025)
por: Wang, Shaowen, et al.
Publicado: (2025)
Exploring the Capability of Mamba in Speech Applications
por: Miyazaki, Koichi, et al.
Publicado: (2024)
por: Miyazaki, Koichi, et al.
Publicado: (2024)
A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement
por: Zhang, Yuewei, et al.
Publicado: (2024)
por: Zhang, Yuewei, et al.
Publicado: (2024)
Ejemplares similares
-
Pretrained Conformers for Audio Fingerprinting and Retrieval
por: Altwlkany, Kemal, et al.
Publicado: (2025) -
A Recurrent Neural Network Approach to the Answering Machine Detection Problem
por: Altwlkany, Kemal, et al.
Publicado: (2024) -
On the Language and Gender Biases in PSTN, VoIP and Neural Audio Codecs
por: Altwlkany, Kemal, et al.
Publicado: (2025) -
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
por: Bhattacharjee, Aditya, et al.
Publicado: (2025) -
PeakNetFP: Peak-based Neural Audio Fingerprinting Robust to Extreme Time Stretching
por: Cortès-Sebastià, Guillem, et al.
Publicado: (2025)