Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Comanducci, Luca, Antonacci, Fabio, Sarti, Augusto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Room Transfer Function Reconstruction Using Complex-valued Neural Networks and Irregularly Distributed Microphones
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
Physics-Informed Transfer Learning for Data-Driven Sound Source Reconstruction in Near-Field Acoustic Holography
von: Luan, Xinmeng, et al.
Veröffentlicht: (2025)
von: Luan, Xinmeng, et al.
Veröffentlicht: (2025)
Synthetic training set generation using text-to-audio models for environmental sound classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
Reconstruction of Sound Field through Diffusion Models
von: Miotello, Federico, et al.
Veröffentlicht: (2023)
von: Miotello, Federico, et al.
Veröffentlicht: (2023)
Synthesis of Soundfields through Irregular Loudspeaker Arrays Based on Convolutional Neural Networks
von: Comanducci, Luca, et al.
Veröffentlicht: (2022)
von: Comanducci, Luca, et al.
Veröffentlicht: (2022)
PAGURI: a user experience study of creative interaction with text-to-music models
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
MambaFoley: Foley Sound Generation using Selective State-Space Models
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024)
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024)
Towards HRTF Personalization using Denoising Diffusion Models
von: Sánchez, Juan Camilo Albarracín, et al.
Veröffentlicht: (2025)
von: Sánchez, Juan Camilo Albarracín, et al.
Veröffentlicht: (2025)
Acoustic source localization in the spherical harmonics domain exploiting low-rank approximations
von: Cobos, Maximo, et al.
Veröffentlicht: (2023)
von: Cobos, Maximo, et al.
Veröffentlicht: (2023)
AI-Assisted Music Production: A User Study on Text-to-Music Models
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
Implicit neural representation with physics-informed neural networks for the reconstruction of the early part of room impulse responses
von: Pezzoli, Mirco, et al.
Veröffentlicht: (2023)
von: Pezzoli, Mirco, et al.
Veröffentlicht: (2023)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
von: Messina, Francisco, et al.
Veröffentlicht: (2025)
von: Messina, Francisco, et al.
Veröffentlicht: (2025)
Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
von: Passoni, Riccardo, et al.
Veröffentlicht: (2025)
von: Passoni, Riccardo, et al.
Veröffentlicht: (2025)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
End-to-End Diarization utilizing Attractor Deep Clustering
von: Palzer, David, et al.
Veröffentlicht: (2025)
von: Palzer, David, et al.
Veröffentlicht: (2025)
AADNet: An End-to-End Deep Learning Model for Auditory Attention Decoding
von: Nguyen, Nhan Duc Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Nhan Duc Thanh, et al.
Veröffentlicht: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
Layer-wise Analysis for Quality of Multilingual Synthesized Speech
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
Low-Rank Adaptation of Deep Prior Neural Networks For Room Impulse Response Reconstruction
von: Pezzoli, Mirco, et al.
Veröffentlicht: (2025)
von: Pezzoli, Mirco, et al.
Veröffentlicht: (2025)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2024)
von: Chen, Jinming, et al.
Veröffentlicht: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Toward Deep Drum Source Separation
von: Mezza, Alessandro Ilic, et al.
Veröffentlicht: (2023)
von: Mezza, Alessandro Ilic, et al.
Veröffentlicht: (2023)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
von: Raimon, Athul, et al.
Veröffentlicht: (2024)
von: Raimon, Athul, et al.
Veröffentlicht: (2024)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
End-to-End Amp Modeling: From Data to Controllable Guitar Amplifier Models
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
A Physics-Informed Neural Network-Based Approach for the Spatial Upsampling of Spherical Microphone Arrays
von: Miotello, Federico, et al.
Veröffentlicht: (2024)
von: Miotello, Federico, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Room Transfer Function Reconstruction Using Complex-valued Neural Networks and Irregularly Distributed Microphones
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024) -
Physics-Informed Transfer Learning for Data-Driven Sound Source Reconstruction in Near-Field Acoustic Holography
von: Luan, Xinmeng, et al.
Veröffentlicht: (2025) -
Synthetic training set generation using text-to-audio models for environmental sound classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024) -
Reconstruction of Sound Field through Diffusion Models
von: Miotello, Federico, et al.
Veröffentlicht: (2023) -
Synthesis of Soundfields through Irregular Loudspeaker Arrays Based on Convolutional Neural Networks
von: Comanducci, Luca, et al.
Veröffentlicht: (2022)