An Explainable Proxy Model for Multiabel Audio Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mariotte, Théo, Almudévar, Antonio, Tahon, Marie, Ortega, Alfonso |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparse Autoencoders Make Audio Foundation Models more Explainable
von: Mariotte, Théo, et al.
Veröffentlicht: (2025)
von: Mariotte, Théo, et al.
Veröffentlicht: (2025)
Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
von: Lebourdais, Martin, et al.
Veröffentlicht: (2024)
von: Lebourdais, Martin, et al.
Veröffentlicht: (2024)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024)
von: Luo, Yi, et al.
Veröffentlicht: (2024)
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
ANIRA: An Architecture for Neural Network Inference in Real-Time Audio Applications
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
Angular Distance Distribution Loss for Audio Classification
von: Almudévar, Antonio, et al.
Veröffentlicht: (2024)
von: Almudévar, Antonio, et al.
Veröffentlicht: (2024)
AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
von: Siddiqui, Md. Saiful Bari, et al.
Veröffentlicht: (2025)
von: Siddiqui, Md. Saiful Bari, et al.
Veröffentlicht: (2025)
Automatic Voice Identification after Speech Resynthesis using PPG
von: Gaudier, Thibault, et al.
Veröffentlicht: (2024)
von: Gaudier, Thibault, et al.
Veröffentlicht: (2024)
AudioGenX: Explainability on Text-to-Audio Generative Models
von: Kang, Hyunju, et al.
Veröffentlicht: (2025)
von: Kang, Hyunju, et al.
Veröffentlicht: (2025)
Model as Loss: A Self-Consistent Training Paradigm
von: Phaye, Saisamarth Rajesh, et al.
Veröffentlicht: (2025)
von: Phaye, Saisamarth Rajesh, et al.
Veröffentlicht: (2025)
MaskSR: Masked Language Model for Full-band Speech Restoration
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
von: Huang, Zhaolan, et al.
Veröffentlicht: (2024)
von: Huang, Zhaolan, et al.
Veröffentlicht: (2024)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
von: Ali, Shams Nafisa, et al.
Veröffentlicht: (2024)
von: Ali, Shams Nafisa, et al.
Veröffentlicht: (2024)
Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
Real-time Timbre Remapping with Differentiable DSP
von: Shier, Jordie, et al.
Veröffentlicht: (2024)
von: Shier, Jordie, et al.
Veröffentlicht: (2024)
Deep Active Speech Cancellation with Mamba-Masking Network
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
Design Of Rubble Analyzer Probe Using ML For Earthquake
von: Sebastian, Abhishek, et al.
Veröffentlicht: (2023)
von: Sebastian, Abhishek, et al.
Veröffentlicht: (2023)
Gaussian Process Regression of Steering Vectors With Physics-Aware Deep Composite Kernels for Augmented Listening
von: Di Carlo, Diego, et al.
Veröffentlicht: (2025)
von: Di Carlo, Diego, et al.
Veröffentlicht: (2025)
Pitch-Conditioned Instrument Sound Synthesis From an Interactive Timbre Latent Space
von: Limberg, Christian, et al.
Veröffentlicht: (2025)
von: Limberg, Christian, et al.
Veröffentlicht: (2025)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
von: Kang, Minsu, et al.
Veröffentlicht: (2025)
von: Kang, Minsu, et al.
Veröffentlicht: (2025)
Compressing Quaternion Convolutional Neural Networks for Audio Classification
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
Listenable Maps for Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
von: Kim, Minje, et al.
Veröffentlicht: (2024)
von: Kim, Minje, et al.
Veröffentlicht: (2024)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
von: Huang, Kuan-Tang, et al.
Veröffentlicht: (2026)
von: Huang, Kuan-Tang, et al.
Veröffentlicht: (2026)
Listenable Maps for Zero-Shot Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
Latent Granular Resynthesis using Neural Audio Codecs
von: Tokui, Nao, et al.
Veröffentlicht: (2025)
von: Tokui, Nao, et al.
Veröffentlicht: (2025)
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
Resampling Filter Design for Multirate Neural Audio Effect Processing
von: Carson, Alistair, et al.
Veröffentlicht: (2025)
von: Carson, Alistair, et al.
Veröffentlicht: (2025)
A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Sparse Autoencoders Make Audio Foundation Models more Explainable
von: Mariotte, Théo, et al.
Veröffentlicht: (2025) -
Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
von: Lebourdais, Martin, et al.
Veröffentlicht: (2024) -
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025) -
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024) -
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)