Gespeichert in:
| Hauptverfasser: | Taylor, James, Mack, Wolfgang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.20036 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prototypical Contrastive Learning For Improved Few-Shot Audio Classification
von: Sgouropoulos, Christos, et al.
Veröffentlicht: (2025)
von: Sgouropoulos, Christos, et al.
Veröffentlicht: (2025)
Leveraging Prediction Entropy for Automatic Prompt Weighting in Zero-Shot Audio-Language Classification
von: Khoury, Karim El, et al.
Veröffentlicht: (2026)
von: Khoury, Karim El, et al.
Veröffentlicht: (2026)
On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
von: Heggan, Calum, et al.
Veröffentlicht: (2024)
von: Heggan, Calum, et al.
Veröffentlicht: (2024)
Multi-label Zero-Shot Audio Classification with Temporal Attention
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
Listenable Maps for Zero-Shot Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
When Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper
von: Islam, Akif, et al.
Veröffentlicht: (2026)
von: Islam, Akif, et al.
Veröffentlicht: (2026)
Self-Supervised Learning for Few-Shot Bird Sound Classification
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
Blind Audio Bandwidth Extension: A Diffusion-Based Zero-Shot Approach
von: Moliner, Eloi, et al.
Veröffentlicht: (2023)
von: Moliner, Eloi, et al.
Veröffentlicht: (2023)
Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion
von: Manor, Hila, et al.
Veröffentlicht: (2024)
von: Manor, Hila, et al.
Veröffentlicht: (2024)
Embedding-Space Diffusion for Zero-Shot Environmental Sound Classification
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
von: Marincione, Davide, et al.
Veröffentlicht: (2025)
von: Marincione, Davide, et al.
Veröffentlicht: (2025)
Exploring Meta Information for Audio-based Zero-shot Bird Classification
von: Gebhard, Alexander, et al.
Veröffentlicht: (2023)
von: Gebhard, Alexander, et al.
Veröffentlicht: (2023)
Towards Robust Few-shot Class Incremental Learning in Audio Classification using Contrastive Representation
von: Singh, Riyansha, et al.
Veröffentlicht: (2024)
von: Singh, Riyansha, et al.
Veröffentlicht: (2024)
APEX: Audio Prototype EXplanations for Classification Tasks
von: Kawa, Piotr, et al.
Veröffentlicht: (2026)
von: Kawa, Piotr, et al.
Veröffentlicht: (2026)
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification
von: Sundar, Anirudh S., et al.
Veröffentlicht: (2023)
von: Sundar, Anirudh S., et al.
Veröffentlicht: (2023)
Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
An Automated Pipeline for Few-Shot Bird Call Classification: A Case Study with the Tooth-Billed Pigeon
von: Jana, Abhishek, et al.
Veröffentlicht: (2025)
von: Jana, Abhishek, et al.
Veröffentlicht: (2025)
Adaptive Discovery of Interpretable Audio Attributes with Multimodal LLMs for Low-Resource Classification
von: Yoshimura, Kosuke, et al.
Veröffentlicht: (2026)
von: Yoshimura, Kosuke, et al.
Veröffentlicht: (2026)
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
Bringing the Discussion of Minima Sharpness to the Audio Domain: a Filter-Normalised Evaluation for Acoustic Scene Classification
von: Milling, Manuel, et al.
Veröffentlicht: (2023)
von: Milling, Manuel, et al.
Veröffentlicht: (2023)
Zero-Shot Mono-to-Binaural Speech Synthesis
von: Levkovitch, Alon, et al.
Veröffentlicht: (2024)
von: Levkovitch, Alon, et al.
Veröffentlicht: (2024)
Low-Resource Guidance for Controllable Latent Audio Diffusion
von: Novack, Zachary, et al.
Veröffentlicht: (2026)
von: Novack, Zachary, et al.
Veröffentlicht: (2026)
Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2025)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2025)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
von: Yang, Mu, et al.
Veröffentlicht: (2024)
von: Yang, Mu, et al.
Veröffentlicht: (2024)
How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection
von: Xiao, Yixuan, et al.
Veröffentlicht: (2026)
von: Xiao, Yixuan, et al.
Veröffentlicht: (2026)
ADNAC: Audio Denoiser using Neural Audio Codec
von: Jimon, Daniel, et al.
Veröffentlicht: (2025)
von: Jimon, Daniel, et al.
Veröffentlicht: (2025)
Improving Out-of-Domain Audio Deepfake Detection via Layer Selection and Fusion of SSL-Based Countermeasures
von: Serrano, Pierre, et al.
Veröffentlicht: (2025)
von: Serrano, Pierre, et al.
Veröffentlicht: (2025)
Pre-Finetuning for Few-Shot Emotional Speech Recognition
von: Chen, Maximillian, et al.
Veröffentlicht: (2023)
von: Chen, Maximillian, et al.
Veröffentlicht: (2023)
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning
von: Wu, Daiqing, et al.
Veröffentlicht: (2026)
von: Wu, Daiqing, et al.
Veröffentlicht: (2026)
Zero-Shot Multi-Lingual Speaker Verification in Clinical Trials
von: Akram, Ali, et al.
Veröffentlicht: (2024)
von: Akram, Ali, et al.
Veröffentlicht: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
von: Janiczek, John, et al.
Veröffentlicht: (2024)
von: Janiczek, John, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Prototypical Contrastive Learning For Improved Few-Shot Audio Classification
von: Sgouropoulos, Christos, et al.
Veröffentlicht: (2025) -
Leveraging Prediction Entropy for Automatic Prompt Weighting in Zero-Shot Audio-Language Classification
von: Khoury, Karim El, et al.
Veröffentlicht: (2026) -
On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
von: Heggan, Calum, et al.
Veröffentlicht: (2024) -
Multi-label Zero-Shot Audio Classification with Temporal Attention
von: Dogan, Duygu, et al.
Veröffentlicht: (2024) -
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
von: Anand, Nishit, et al.
Veröffentlicht: (2024)