Focal Modulation Networks for Interpretable Sound Classification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Della Libera, Luca, Subakan, Cem, Ravanelli, Mirco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
Autoregressive Speech Enhancement via Acoustic Tokens
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
Listenable Maps for Zero-Shot Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
LL-SDR: Low-Latency Speech enhancement through Discrete Representations
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
Audio Editing with Non-Rigid Text Prompts
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
Resource-Efficient Separation Transformer
von: Della Libera, Luca, et al.
Veröffentlicht: (2022)
von: Della Libera, Luca, et al.
Veröffentlicht: (2022)
Listenable Maps for Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
Toward Faithful Explanations in Acoustic Anomaly Detection
von: Elrashid, Maab, et al.
Veröffentlicht: (2026)
von: Elrashid, Maab, et al.
Veröffentlicht: (2026)
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
DASB - Discrete Audio and Speech Benchmark
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
Dynamic HumTrans: Humming Transcription Using CNNs and Dynamic Programming
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
Investigating Faithfulness in Large Audio Language Models
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
Beyond Fixed Frames: Dynamic Character-Aligned Speech Tokenization
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
Spectrotemporal Modulation: Efficient and Interpretable Feature Representation for Classifying Speech, Music, and Environmental Sounds
von: Chang, Andrew, et al.
Veröffentlicht: (2025)
von: Chang, Andrew, et al.
Veröffentlicht: (2025)
Reconstruction of Sound Field through Diffusion Models
von: Miotello, Federico, et al.
Veröffentlicht: (2023)
von: Miotello, Federico, et al.
Veröffentlicht: (2023)
Feature Aggregation in Joint Sound Classification and Localization Neural Networks
von: Healy, Brendan, et al.
Veröffentlicht: (2023)
von: Healy, Brendan, et al.
Veröffentlicht: (2023)
Classification of Short Segment Pediatric Heart Sounds Based on a Transformer-Based Convolutional Neural Network
von: Hassanuzzaman, Md, et al.
Veröffentlicht: (2024)
von: Hassanuzzaman, Md, et al.
Veröffentlicht: (2024)
Exploring Token-Space Manipulation in Latent Audio Tokenizers
von: Paissan, Francesco, et al.
Veröffentlicht: (2026)
von: Paissan, Francesco, et al.
Veröffentlicht: (2026)
Advanced Framework for Animal Sound Classification With Features Optimization
von: Yang, Qiang, et al.
Veröffentlicht: (2024)
von: Yang, Qiang, et al.
Veröffentlicht: (2024)
Room Transfer Function Reconstruction Using Complex-valued Neural Networks and Irregularly Distributed Microphones
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
Mixture of Mixups for Multi-label Classification of Rare Anuran Sounds
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
Embedding-Space Diffusion for Zero-Shot Environmental Sound Classification
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
Self-Supervised Learning for Few-Shot Bird Sound Classification
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
Lungmix: A Mixup-Based Strategy for Generalization in Respiratory Sound Classification
von: Ge, Shijia, et al.
Veröffentlicht: (2024)
von: Ge, Shijia, et al.
Veröffentlicht: (2024)
Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
von: Bae, Sangmin, et al.
Veröffentlicht: (2023)
von: Bae, Sangmin, et al.
Veröffentlicht: (2023)
Patient-Aware Feature Alignment for Robust Lung Sound Classification:Cohesion-Separation and Global Alignment Losses
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2025)
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2025)
Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
von: Park, Sangwook, et al.
Veröffentlicht: (2021)
von: Park, Sangwook, et al.
Veröffentlicht: (2021)
SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
ProGRes: Prompted Generative Rescoring on ASR n-Best
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
von: Della Libera, Luca, et al.
Veröffentlicht: (2025) -
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2025) -
Autoregressive Speech Enhancement via Acoustic Tokens
von: Della Libera, Luca, et al.
Veröffentlicht: (2025) -
Listenable Maps for Zero-Shot Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024) -
LL-SDR: Low-Latency Speech enhancement through Discrete Representations
von: Li, Jingyi, et al.
Veröffentlicht: (2026)