Speech Loudness in Broadcasting and Streaming
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Torcoli, Matteo, Halimeh, Mhd Modar, Leitz, Thomas, Grewe, Yannik, Kratschmer, Michael, Neugebauer, Bernhard, Murtaza, Adrian, Fuchs, Harald, Habets, Emanuël A. P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
von: Halimeh, Mhd Modar, et al.
Veröffentlicht: (2024)
von: Halimeh, Mhd Modar, et al.
Veröffentlicht: (2024)
Navigating PESQ: Up-to-Date Versions and Open Implementations
von: Torcoli, Matteo, et al.
Veröffentlicht: (2025)
von: Torcoli, Matteo, et al.
Veröffentlicht: (2025)
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
von: Halimeh, Mhd Modar, et al.
Veröffentlicht: (2025)
von: Halimeh, Mhd Modar, et al.
Veröffentlicht: (2025)
Robust Speech Activity Detection in the Presence of Singing Voice
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
ODAQ: Open Dataset of Audio Quality
von: Torcoli, Matteo, et al.
Veröffentlicht: (2023)
von: Torcoli, Matteo, et al.
Veröffentlicht: (2023)
Neural Directional Filtering Using a Compact Microphone Array
von: Huang, Weilong, et al.
Veröffentlicht: (2025)
von: Huang, Weilong, et al.
Veröffentlicht: (2025)
Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array
von: Wechsler, Julian, et al.
Veröffentlicht: (2024)
von: Wechsler, Julian, et al.
Veröffentlicht: (2024)
Expanding and Analyzing ODAQ -- the Open Dataset of Audio Quality
von: Dick, Sascha, et al.
Veröffentlicht: (2025)
von: Dick, Sascha, et al.
Veröffentlicht: (2025)
Dynamic Slimmable Networks for Efficient Speech Separation
von: Elminshawi, Mohamed, et al.
Veröffentlicht: (2025)
von: Elminshawi, Mohamed, et al.
Veröffentlicht: (2025)
Benchmarking Neural Speech Codec Intelligibility with SITool
von: Leschanowsky, Anna, et al.
Veröffentlicht: (2025)
von: Leschanowsky, Anna, et al.
Veröffentlicht: (2025)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
von: Lakshminarayana, Kishor Kayyar, et al.
Veröffentlicht: (2025)
von: Lakshminarayana, Kishor Kayyar, et al.
Veröffentlicht: (2025)
Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks
von: Götz, Philipp, et al.
Veröffentlicht: (2025)
von: Götz, Philipp, et al.
Veröffentlicht: (2025)
Sample Rate Offset Compensated Acoustic Echo Cancellation For Multi-Device Scenarios
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Neural Directional Filtering with Configurable Directivity Pattern at Inference
von: Huang, Weilong, et al.
Veröffentlicht: (2025)
von: Huang, Weilong, et al.
Veröffentlicht: (2025)
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
Stereo Reproduction in the Presence of Sample Rate Offsets
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
VoxATtack: A Multimodal Attack on Voice Anonymization Systems
von: Aloradi, Ahmad, et al.
Veröffentlicht: (2025)
von: Aloradi, Ahmad, et al.
Veröffentlicht: (2025)
Room Impulse Response Completion Using Signal-Prediction Diffusion Models Conditioned on Simulated Early Reflections
von: Xu, Zeyu, et al.
Veröffentlicht: (2026)
von: Xu, Zeyu, et al.
Veröffentlicht: (2026)
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
Comparative Analysis Of Discriminative Deep Learning-Based Noise Reduction Methods In Low SNR Scenarios
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Investigating the impact of stereo processing -- a study for extending the Open Dataset of Audio Quality (ODAQ)
von: Dick, Sascha, et al.
Veröffentlicht: (2025)
von: Dick, Sascha, et al.
Veröffentlicht: (2025)
Blind Acoustic Parameter Estimation Through Task-Agnostic Embeddings Using Latent Approximations
von: Götz, Philipp, et al.
Veröffentlicht: (2024)
von: Götz, Philipp, et al.
Veröffentlicht: (2024)
NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction
von: Huang, Weilong, et al.
Veröffentlicht: (2026)
von: Huang, Weilong, et al.
Veröffentlicht: (2026)
Data-driven Joint Detection and Localization of Acoustic Reflectors
von: Bicer, H. Nazim, et al.
Veröffentlicht: (2024)
von: Bicer, H. Nazim, et al.
Veröffentlicht: (2024)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
spINAch: A Diachronic Corpus of French Broadcast Speech Controlled for Speakers' Age and Gender
von: Devauchelle, Simon, et al.
Veröffentlicht: (2026)
von: Devauchelle, Simon, et al.
Veröffentlicht: (2026)
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
von: Poncelet, Jakob, et al.
Veröffentlicht: (2025)
von: Poncelet, Jakob, et al.
Veröffentlicht: (2025)
Chunkwise Aligners for Streaming Speech Recognition
von: Teo, Wen Shen, et al.
Veröffentlicht: (2026)
von: Teo, Wen Shen, et al.
Veröffentlicht: (2026)
Assessing the Impact of Noise and Speech Enhancement on the Intelligibility of Speech Codecs
von: Behringer, Lyonel, et al.
Veröffentlicht: (2026)
von: Behringer, Lyonel, et al.
Veröffentlicht: (2026)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Large Model Empowered Streaming Speech Semantic Communications
von: Weng, Zhenzi, et al.
Veröffentlicht: (2025)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2025)
A Hybrid Approach for Low-Complexity Joint Acoustic Echo and Noise Reduction
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
FastEnhancer: Speed-Optimized Streaming Neural Speech Enhancement
von: Ahn, Sunghwan, et al.
Veröffentlicht: (2025)
von: Ahn, Sunghwan, et al.
Veröffentlicht: (2025)
Align-ULCNet: Towards Low-Complexity and Robust Acoustic Echo and Noise Reduction
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
von: Shukla, Sakshi Deo, et al.
Veröffentlicht: (2024)
von: Shukla, Sakshi Deo, et al.
Veröffentlicht: (2024)
Streaming Speech-to-Confusion Network Speech Recognition
von: Filimonov, Denis, et al.
Veröffentlicht: (2023)
von: Filimonov, Denis, et al.
Veröffentlicht: (2023)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
von: Halimeh, Mhd Modar, et al.
Veröffentlicht: (2024) -
Navigating PESQ: Up-to-Date Versions and Open Implementations
von: Torcoli, Matteo, et al.
Veröffentlicht: (2025) -
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
von: Halimeh, Mhd Modar, et al.
Veröffentlicht: (2025) -
Robust Speech Activity Detection in the Presence of Singing Voice
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025) -
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)