Multi-label Zero-Shot Audio Classification with Temporal Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dogan, Duygu, Xie, Huang, Heittola, Toni, Virtanen, Tuomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A decade of DCASE: Achievements, practices, evaluations and future challenges
von: Mesaros, Annamaria, et al.
Veröffentlicht: (2024)
von: Mesaros, Annamaria, et al.
Veröffentlicht: (2024)
Multi-Channel Replay Speech Detection using Acoustic Maps
von: Neri, Michael, et al.
Veröffentlicht: (2026)
von: Neri, Michael, et al.
Veröffentlicht: (2026)
Automatic Contextual Audio Denoising
von: Luong, Diep, et al.
Veröffentlicht: (2026)
von: Luong, Diep, et al.
Veröffentlicht: (2026)
Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
von: Moritz, Martin, et al.
Veröffentlicht: (2024)
von: Moritz, Martin, et al.
Veröffentlicht: (2024)
Acoustic Scene Classification: A Competition Review
von: Gharib, Shayan, et al.
Veröffentlicht: (2018)
von: Gharib, Shayan, et al.
Veröffentlicht: (2018)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
From Weak to Strong Sound Event Labels using Adaptive Change-Point Detection and Active Learning
von: Martinsson, John, et al.
Veröffentlicht: (2024)
von: Martinsson, John, et al.
Veröffentlicht: (2024)
Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2025)
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2025)
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
von: Luong, Diep, et al.
Veröffentlicht: (2025)
von: Luong, Diep, et al.
Veröffentlicht: (2025)
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
von: Zhang, Shiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Shiqi, et al.
Veröffentlicht: (2025)
Adversarial Representation Learning for Robust Privacy Preservation in Audio
von: Gharib, Shayan, et al.
Veröffentlicht: (2023)
von: Gharib, Shayan, et al.
Veröffentlicht: (2023)
Score-informed Music Source Separation: Improving Synthetic-to-real Generalization in Classical Music
von: Tunturi, Eetu, et al.
Veröffentlicht: (2025)
von: Tunturi, Eetu, et al.
Veröffentlicht: (2025)
On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
Representation Learning for Audio Privacy Preservation using Source Separation and Robust Adversarial Learning
von: Luong, Diep, et al.
Veröffentlicht: (2023)
von: Luong, Diep, et al.
Veröffentlicht: (2023)
Listenable Maps for Zero-Shot Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion
von: Manor, Hila, et al.
Veröffentlicht: (2024)
von: Manor, Hila, et al.
Veröffentlicht: (2024)
Blind Audio Bandwidth Extension: A Diffusion-Based Zero-Shot Approach
von: Moliner, Eloi, et al.
Veröffentlicht: (2023)
von: Moliner, Eloi, et al.
Veröffentlicht: (2023)
On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
von: Heggan, Calum, et al.
Veröffentlicht: (2024)
von: Heggan, Calum, et al.
Veröffentlicht: (2024)
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier
von: Dumpis, Martynas, et al.
Veröffentlicht: (2026)
von: Dumpis, Martynas, et al.
Veröffentlicht: (2026)
Embedding-Space Diffusion for Zero-Shot Environmental Sound Classification
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
Exploring Meta Information for Audio-based Zero-shot Bird Classification
von: Gebhard, Alexander, et al.
Veröffentlicht: (2023)
von: Gebhard, Alexander, et al.
Veröffentlicht: (2023)
Mixture of Mixups for Multi-label Classification of Rare Anuran Sounds
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification
von: Sundar, Anirudh S., et al.
Veröffentlicht: (2023)
von: Sundar, Anirudh S., et al.
Veröffentlicht: (2023)
Low-Complexity Acoustic Scene Classification with Device Information in the DCASE 2025 Challenge
von: Schmid, Florian, et al.
Veröffentlicht: (2025)
von: Schmid, Florian, et al.
Veröffentlicht: (2025)
Zero-Shot Multi-Lingual Speaker Verification in Clinical Trials
von: Akram, Ali, et al.
Veröffentlicht: (2024)
von: Akram, Ali, et al.
Veröffentlicht: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
von: Janiczek, John, et al.
Veröffentlicht: (2024)
von: Janiczek, John, et al.
Veröffentlicht: (2024)
Automatic Live Music Song Identification Using Multi-level Deep Sequence Similarity Learning
von: Hakala, Aapo, et al.
Veröffentlicht: (2025)
von: Hakala, Aapo, et al.
Veröffentlicht: (2025)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
Speaker Distance Estimation in Enclosures from Single-Channel Audio
von: Neri, Michael, et al.
Veröffentlicht: (2024)
von: Neri, Michael, et al.
Veröffentlicht: (2024)
LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
von: Primus, Paul, et al.
Veröffentlicht: (2025)
von: Primus, Paul, et al.
Veröffentlicht: (2025)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
Data-Efficient Low-Complexity Acoustic Scene Classification in the DCASE 2024 Challenge
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation
von: Garcia-Martinez, Jaime, et al.
Veröffentlicht: (2024)
von: Garcia-Martinez, Jaime, et al.
Veröffentlicht: (2024)
Zero-Shot Mono-to-Binaural Speech Synthesis
von: Levkovitch, Alon, et al.
Veröffentlicht: (2024)
von: Levkovitch, Alon, et al.
Veröffentlicht: (2024)
Permutation Invariant Recurrent Neural Networks for Sound Source Tracking Applications
von: Diaz-Guerra, David, et al.
Veröffentlicht: (2023)
von: Diaz-Guerra, David, et al.
Veröffentlicht: (2023)
Multi-Utterance Speech Separation and Association Trained on Short Segments
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
Neural Ambisonics encoding for compact irregular microphone arrays
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2024)
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A decade of DCASE: Achievements, practices, evaluations and future challenges
von: Mesaros, Annamaria, et al.
Veröffentlicht: (2024) -
Multi-Channel Replay Speech Detection using Acoustic Maps
von: Neri, Michael, et al.
Veröffentlicht: (2026) -
Automatic Contextual Audio Denoising
von: Luong, Diep, et al.
Veröffentlicht: (2026) -
Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
von: Moritz, Martin, et al.
Veröffentlicht: (2024) -
Acoustic Scene Classification: A Competition Review
von: Gharib, Shayan, et al.
Veröffentlicht: (2018)