LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mancini, Eleonora, Paissan, Francesco, Ravanelli, Mirco, Subakan, Cem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Listenable Maps for Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
Listenable Maps for Zero-Shot Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
Audio Editing with Non-Rigid Text Prompts
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
Resource-Efficient Separation Transformer
von: Della Libera, Luca, et al.
Veröffentlicht: (2022)
von: Della Libera, Luca, et al.
Veröffentlicht: (2022)
Autoregressive Speech Enhancement via Acoustic Tokens
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
Focal Modulation Networks for Interpretable Sound Classification
von: Della Libera, Luca, et al.
Veröffentlicht: (2024)
von: Della Libera, Luca, et al.
Veröffentlicht: (2024)
Toward Faithful Explanations in Acoustic Anomaly Detection
von: Elrashid, Maab, et al.
Veröffentlicht: (2026)
von: Elrashid, Maab, et al.
Veröffentlicht: (2026)
Dynamic HumTrans: Humming Transcription Using CNNs and Dynamic Programming
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
DASB - Discrete Audio and Speech Benchmark
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
LL-SDR: Low-Latency Speech enhancement through Discrete Representations
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
Investigating Faithfulness in Large Audio Language Models
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
ANIRA: An Architecture for Neural Network Inference in Real-Time Audio Applications
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
An Explainable Proxy Model for Multiabel Audio Segmentation
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024)
von: Luo, Yi, et al.
Veröffentlicht: (2024)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
von: Mancini, Eleonora, et al.
Veröffentlicht: (2025)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2025)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
von: Siddiqui, Md. Saiful Bari, et al.
Veröffentlicht: (2025)
von: Siddiqui, Md. Saiful Bari, et al.
Veröffentlicht: (2025)
T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
Implicit neural representation with physics-informed neural networks for the reconstruction of the early part of room impulse responses
von: Pezzoli, Mirco, et al.
Veröffentlicht: (2023)
von: Pezzoli, Mirco, et al.
Veröffentlicht: (2023)
Room Transfer Function Reconstruction Using Complex-valued Neural Networks and Irregularly Distributed Microphones
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
tinyCLAP: Distilling Constrastive Language-Audio Pretrained Models
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
Deep Active Speech Cancellation with Mamba-Masking Network
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
Design Of Rubble Analyzer Probe Using ML For Earthquake
von: Sebastian, Abhishek, et al.
Veröffentlicht: (2023)
von: Sebastian, Abhishek, et al.
Veröffentlicht: (2023)
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
von: Ali, Shams Nafisa, et al.
Veröffentlicht: (2024)
von: Ali, Shams Nafisa, et al.
Veröffentlicht: (2024)
Gaussian Process Regression of Steering Vectors With Physics-Aware Deep Composite Kernels for Augmented Listening
von: Di Carlo, Diego, et al.
Veröffentlicht: (2025)
von: Di Carlo, Diego, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Listenable Maps for Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024) -
Listenable Maps for Zero-Shot Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024) -
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024) -
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
von: Della Libera, Luca, et al.
Veröffentlicht: (2025) -
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)