AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
Fuente:
arXiv
Saved in:
| Main Authors: | Dutta, Sagar, Arora, Vipul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention-Based Audio Embeddings for Query-by-Example
by: Singh, Anup, et al.
Published: (2022)
by: Singh, Anup, et al.
Published: (2022)
Text-based Audio Retrieval by Learning from Similarities between Audio Captions
by: Xie, Huang, et al.
Published: (2024)
by: Xie, Huang, et al.
Published: (2024)
SyncNet: correlating objective for time delay estimation in audio signals
by: Raina, Akshay, et al.
Published: (2022)
by: Raina, Akshay, et al.
Published: (2022)
Weakly Supervised Tabla Stroke Transcription via TI-SDRM: A Rhythm-Aware Lattice Rescoring Framework
by: Kodag, Rahul Bapusaheb, et al.
Published: (2026)
by: Kodag, Rahul Bapusaheb, et al.
Published: (2026)
Zimtohrli: An Efficient Psychoacoustic Audio Similarity Metric
by: Alakuijala, Jyrki, et al.
Published: (2025)
by: Alakuijala, Jyrki, et al.
Published: (2025)
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models
by: Li, Longhao, et al.
Published: (2026)
by: Li, Longhao, et al.
Published: (2026)
Explainable Deep Learning Analysis for Raga Identification in Indian Art Music
by: Singh, Parampreet, et al.
Published: (2024)
by: Singh, Parampreet, et al.
Published: (2024)
Uncertainty Quantification in Melody Estimation using Histogram Representation
by: Saxena, Kavya Ranjan, et al.
Published: (2025)
by: Saxena, Kavya Ranjan, et al.
Published: (2025)
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
by: Kodag, Rahul Bapusaheb, et al.
Published: (2025)
by: Kodag, Rahul Bapusaheb, et al.
Published: (2025)
Unified Audio Event Detection
by: Jiang, Yidi, et al.
Published: (2024)
by: Jiang, Yidi, et al.
Published: (2024)
Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection
by: Sudan, Jaskirat, et al.
Published: (2026)
by: Sudan, Jaskirat, et al.
Published: (2026)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
by: Kishi, Minoru, et al.
Published: (2025)
by: Kishi, Minoru, et al.
Published: (2025)
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
by: Lin, Jingru, et al.
Published: (2026)
by: Lin, Jingru, et al.
Published: (2026)
TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR
by: Ravi, Nagarathna, et al.
Published: (2024)
by: Ravi, Nagarathna, et al.
Published: (2024)
Audio-Image Cross-Modal Retrieval with Onomatopoeic Images
by: Imoto, Keisuke, et al.
Published: (2026)
by: Imoto, Keisuke, et al.
Published: (2026)
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
by: Singh, Anup, et al.
Published: (2025)
by: Singh, Anup, et al.
Published: (2025)
Improving Active Learning for Melody Estimation by Disentangling Uncertainties
by: Jaiswal, Aayush, et al.
Published: (2025)
by: Jaiswal, Aayush, et al.
Published: (2025)
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
by: Tian, Haokun, et al.
Published: (2025)
by: Tian, Haokun, et al.
Published: (2025)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
by: Dutta, Soumya, et al.
Published: (2024)
by: Dutta, Soumya, et al.
Published: (2024)
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
by: Tsubaki, Shunsuke, et al.
Published: (2024)
by: Tsubaki, Shunsuke, et al.
Published: (2024)
AudioSpa: Spatializing Sound Events with Text
by: Feng, Linfeng, et al.
Published: (2025)
by: Feng, Linfeng, et al.
Published: (2025)
Towards Weakly Supervised Text-to-Audio Grounding
by: Xu, Xuenan, et al.
Published: (2024)
by: Xu, Xuenan, et al.
Published: (2024)
Interactive singing melody extraction based on active adaptation
by: Saxena, Kavya Ranjan, et al.
Published: (2024)
by: Saxena, Kavya Ranjan, et al.
Published: (2024)
$T\bar{a}laGen:$ A System for Automatic $T\bar{a}la$ Identification and Generation
by: Kodag, Rahul Bapusaheb, et al.
Published: (2024)
by: Kodag, Rahul Bapusaheb, et al.
Published: (2024)
InfiniteAudio: Infinite-Length Audio Generation with Consistency
by: Jung, Chaeyoung, et al.
Published: (2025)
by: Jung, Chaeyoung, et al.
Published: (2025)
Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
by: He, Haolin, et al.
Published: (2025)
by: He, Haolin, et al.
Published: (2025)
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
Diff-VS: Efficient Audio-Aware Diffusion U-Net for Vocals Separation
by: Yun-Ning, et al.
Published: (2026)
by: Yun-Ning, et al.
Published: (2026)
Audio-Language Datasets of Scenes and Events: A Survey
by: Wijngaard, Gijs, et al.
Published: (2024)
by: Wijngaard, Gijs, et al.
Published: (2024)
Natural Language Supervision for General-Purpose Audio Representations
by: Elizalde, Benjamin, et al.
Published: (2023)
by: Elizalde, Benjamin, et al.
Published: (2023)
Aud-Sur: An Audio Analyzer Assistant for Audio Surveillance Applications
by: Lam, Phat, et al.
Published: (2025)
by: Lam, Phat, et al.
Published: (2025)
SALAD-VAE: Semantic Audio Compression with Language-Audio Distillation
by: Braun, Sebastian, et al.
Published: (2025)
by: Braun, Sebastian, et al.
Published: (2025)
H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
by: Singh, Akanksha, et al.
Published: (2025)
by: Singh, Akanksha, et al.
Published: (2025)
Identification and Clustering of Unseen Ragas in Indian Art Music
by: Singh, Parampreet, et al.
Published: (2024)
by: Singh, Parampreet, et al.
Published: (2024)
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
by: Bai, Jisheng, et al.
Published: (2024)
by: Bai, Jisheng, et al.
Published: (2024)
Audio Deepfake Detection with Self-Supervised WavLM and Multi-Fusion Attentive Classifier
by: Guo, Yinlin, et al.
Published: (2023)
by: Guo, Yinlin, et al.
Published: (2023)
Data Selection Effects on Self-Supervised Learning of Audio Representations for French Audiovisual Broadcasts
by: Pelloin, Valentin, et al.
Published: (2026)
by: Pelloin, Valentin, et al.
Published: (2026)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
by: Wu, Shih-Lun, et al.
Published: (2023)
by: Wu, Shih-Lun, et al.
Published: (2023)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
by: Yang, Dongchao, et al.
Published: (2023)
by: Yang, Dongchao, et al.
Published: (2023)
BINAQUAL: A Full-Reference Objective Localization Similarity Metric for Binaural Audio
by: Panah, Davoud Shariat, et al.
Published: (2025)
by: Panah, Davoud Shariat, et al.
Published: (2025)
Similar Items
-
Attention-Based Audio Embeddings for Query-by-Example
by: Singh, Anup, et al.
Published: (2022) -
Text-based Audio Retrieval by Learning from Similarities between Audio Captions
by: Xie, Huang, et al.
Published: (2024) -
SyncNet: correlating objective for time delay estimation in audio signals
by: Raina, Akshay, et al.
Published: (2022) -
Weakly Supervised Tabla Stroke Transcription via TI-SDRM: A Rhythm-Aware Lattice Rescoring Framework
by: Kodag, Rahul Bapusaheb, et al.
Published: (2026) -
Zimtohrli: An Efficient Psychoacoustic Audio Similarity Metric
by: Alakuijala, Jyrki, et al.
Published: (2025)