Salvato in:
| Autori principali: | Torcoli, Matteo, Halimeh, Mhd Modar, Leitz, Thomas, Grewe, Yannik, Kratschmer, Michael, Neugebauer, Bernhard, Murtaza, Adrian, Fuchs, Harald, Habets, Emanuël A. P. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2405.17364 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2024)
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2024)
Navigating PESQ: Up-to-Date Versions and Open Implementations
di: Torcoli, Matteo, et al.
Pubblicazione: (2025)
di: Torcoli, Matteo, et al.
Pubblicazione: (2025)
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2025)
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2025)
Robust Speech Activity Detection in the Presence of Singing Voice
di: Grundhuber, Philipp, et al.
Pubblicazione: (2025)
di: Grundhuber, Philipp, et al.
Pubblicazione: (2025)
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
di: Grundhuber, Philipp, et al.
Pubblicazione: (2025)
di: Grundhuber, Philipp, et al.
Pubblicazione: (2025)
ODAQ: Open Dataset of Audio Quality
di: Torcoli, Matteo, et al.
Pubblicazione: (2023)
di: Torcoli, Matteo, et al.
Pubblicazione: (2023)
Neural Directional Filtering Using a Compact Microphone Array
di: Huang, Weilong, et al.
Pubblicazione: (2025)
di: Huang, Weilong, et al.
Pubblicazione: (2025)
Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array
di: Wechsler, Julian, et al.
Pubblicazione: (2024)
di: Wechsler, Julian, et al.
Pubblicazione: (2024)
Expanding and Analyzing ODAQ -- the Open Dataset of Audio Quality
di: Dick, Sascha, et al.
Pubblicazione: (2025)
di: Dick, Sascha, et al.
Pubblicazione: (2025)
Dynamic Slimmable Networks for Efficient Speech Separation
di: Elminshawi, Mohamed, et al.
Pubblicazione: (2025)
di: Elminshawi, Mohamed, et al.
Pubblicazione: (2025)
Benchmarking Neural Speech Codec Intelligibility with SITool
di: Leschanowsky, Anna, et al.
Pubblicazione: (2025)
di: Leschanowsky, Anna, et al.
Pubblicazione: (2025)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2025)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2025)
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
di: Lakshminarayana, Kishor Kayyar, et al.
Pubblicazione: (2025)
di: Lakshminarayana, Kishor Kayyar, et al.
Pubblicazione: (2025)
Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks
di: Götz, Philipp, et al.
Pubblicazione: (2025)
di: Götz, Philipp, et al.
Pubblicazione: (2025)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
Sample Rate Offset Compensated Acoustic Echo Cancellation For Multi-Device Scenarios
di: Korse, Srikanth, et al.
Pubblicazione: (2025)
di: Korse, Srikanth, et al.
Pubblicazione: (2025)
Neural Directional Filtering with Configurable Directivity Pattern at Inference
di: Huang, Weilong, et al.
Pubblicazione: (2025)
di: Huang, Weilong, et al.
Pubblicazione: (2025)
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2025)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2025)
Stereo Reproduction in the Presence of Sample Rate Offsets
di: Korse, Srikanth, et al.
Pubblicazione: (2025)
di: Korse, Srikanth, et al.
Pubblicazione: (2025)
Comparative Analysis Of Discriminative Deep Learning-Based Noise Reduction Methods In Low SNR Scenarios
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
VoxATtack: A Multimodal Attack on Voice Anonymization Systems
di: Aloradi, Ahmad, et al.
Pubblicazione: (2025)
di: Aloradi, Ahmad, et al.
Pubblicazione: (2025)
Room Impulse Response Completion Using Signal-Prediction Diffusion Models Conditioned on Simulated Early Reflections
di: Xu, Zeyu, et al.
Pubblicazione: (2026)
di: Xu, Zeyu, et al.
Pubblicazione: (2026)
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
di: Korse, Srikanth, et al.
Pubblicazione: (2025)
di: Korse, Srikanth, et al.
Pubblicazione: (2025)
Blind Acoustic Parameter Estimation Through Task-Agnostic Embeddings Using Latent Approximations
di: Götz, Philipp, et al.
Pubblicazione: (2024)
di: Götz, Philipp, et al.
Pubblicazione: (2024)
NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction
di: Huang, Weilong, et al.
Pubblicazione: (2026)
di: Huang, Weilong, et al.
Pubblicazione: (2026)
Investigating the impact of stereo processing -- a study for extending the Open Dataset of Audio Quality (ODAQ)
di: Dick, Sascha, et al.
Pubblicazione: (2025)
di: Dick, Sascha, et al.
Pubblicazione: (2025)
Data-driven Joint Detection and Localization of Acoustic Reflectors
di: Bicer, H. Nazim, et al.
Pubblicazione: (2024)
di: Bicer, H. Nazim, et al.
Pubblicazione: (2024)
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
di: Poncelet, Jakob, et al.
Pubblicazione: (2025)
di: Poncelet, Jakob, et al.
Pubblicazione: (2025)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
A Hybrid Approach for Low-Complexity Joint Acoustic Echo and Noise Reduction
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
Align-ULCNet: Towards Low-Complexity and Robust Acoustic Echo and Noise Reduction
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
spINAch: A Diachronic Corpus of French Broadcast Speech Controlled for Speakers' Age and Gender
di: Devauchelle, Simon, et al.
Pubblicazione: (2026)
di: Devauchelle, Simon, et al.
Pubblicazione: (2026)
Assessing the Impact of Noise and Speech Enhancement on the Intelligibility of Speech Codecs
di: Behringer, Lyonel, et al.
Pubblicazione: (2026)
di: Behringer, Lyonel, et al.
Pubblicazione: (2026)
Chunkwise Aligners for Streaming Speech Recognition
di: Teo, Wen Shen, et al.
Pubblicazione: (2026)
di: Teo, Wen Shen, et al.
Pubblicazione: (2026)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
di: Lux, Florian, et al.
Pubblicazione: (2024)
di: Lux, Florian, et al.
Pubblicazione: (2024)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
di: Yang, Yifan, et al.
Pubblicazione: (2024)
di: Yang, Yifan, et al.
Pubblicazione: (2024)
Low-Complexity Neural Wind Noise Reduction for Audio Recordings
di: Eftekhari, Hesam, et al.
Pubblicazione: (2025)
di: Eftekhari, Hesam, et al.
Pubblicazione: (2025)
Large Model Empowered Streaming Speech Semantic Communications
di: Weng, Zhenzi, et al.
Pubblicazione: (2025)
di: Weng, Zhenzi, et al.
Pubblicazione: (2025)
Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
di: Shukla, Sakshi Deo, et al.
Pubblicazione: (2024)
di: Shukla, Sakshi Deo, et al.
Pubblicazione: (2024)
You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks
di: Gaznepoglu, Ünal Ege, et al.
Pubblicazione: (2025)
di: Gaznepoglu, Ünal Ege, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2024) -
Navigating PESQ: Up-to-Date Versions and Open Implementations
di: Torcoli, Matteo, et al.
Pubblicazione: (2025) -
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
di: Halimeh, Mhd Modar, et al.
Pubblicazione: (2025) -
Robust Speech Activity Detection in the Presence of Singing Voice
di: Grundhuber, Philipp, et al.
Pubblicazione: (2025) -
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
di: Grundhuber, Philipp, et al.
Pubblicazione: (2025)