Benchmarking Neural Speech Codec Intelligibility with SITool
Fuente:
arXiv
Saved in:
| Main Authors: | Leschanowsky, Anna, Lakshminarayana, Kishor Kayyar, Rajasekhar, Anjana, Behringer, Lyonel, Kilinc, Ibrahim, Fuchs, Guillaume, Habets, Emanuël A. P. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing the Impact of Noise and Speech Enhancement on the Intelligibility of Speech Codecs
by: Behringer, Lyonel, et al.
Published: (2026)
by: Behringer, Lyonel, et al.
Published: (2026)
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
by: Lakshminarayana, Kishor Kayyar, et al.
Published: (2025)
by: Lakshminarayana, Kishor Kayyar, et al.
Published: (2025)
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
by: Brendel, Andreas, et al.
Published: (2024)
by: Brendel, Andreas, et al.
Published: (2024)
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
by: Halimeh, Mhd Modar, et al.
Published: (2025)
by: Halimeh, Mhd Modar, et al.
Published: (2025)
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
by: Grundhuber, Philipp, et al.
Published: (2025)
by: Grundhuber, Philipp, et al.
Published: (2025)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
by: Lux, Florian, et al.
Published: (2024)
by: Lux, Florian, et al.
Published: (2024)
Examining the Interplay Between Privacy and Fairness for Speech Processing: A Review and Perspective
by: Leschanowsky, Anna, et al.
Published: (2024)
by: Leschanowsky, Anna, et al.
Published: (2024)
Dynamic Slimmable Networks for Efficient Speech Separation
by: Elminshawi, Mohamed, et al.
Published: (2025)
by: Elminshawi, Mohamed, et al.
Published: (2025)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
by: Shetu, Shrishti Saha, et al.
Published: (2025)
by: Shetu, Shrishti Saha, et al.
Published: (2025)
Neural Directional Filtering with Configurable Directivity Pattern at Inference
by: Huang, Weilong, et al.
Published: (2025)
by: Huang, Weilong, et al.
Published: (2025)
Robust Speech Activity Detection in the Presence of Singing Voice
by: Grundhuber, Philipp, et al.
Published: (2025)
by: Grundhuber, Philipp, et al.
Published: (2025)
You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks
by: Gaznepoglu, Ünal Ege, et al.
Published: (2025)
by: Gaznepoglu, Ünal Ege, et al.
Published: (2025)
Speech Loudness in Broadcasting and Streaming
by: Torcoli, Matteo, et al.
Published: (2024)
by: Torcoli, Matteo, et al.
Published: (2024)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
by: Shetu, Shrishti Saha, et al.
Published: (2024)
by: Shetu, Shrishti Saha, et al.
Published: (2024)
Personalized Neural Speech Codec
by: Jang, Inseon, et al.
Published: (2024)
by: Jang, Inseon, et al.
Published: (2024)
Sample Rate Offset Compensated Acoustic Echo Cancellation For Multi-Device Scenarios
by: Korse, Srikanth, et al.
Published: (2025)
by: Korse, Srikanth, et al.
Published: (2025)
NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction
by: Huang, Weilong, et al.
Published: (2026)
by: Huang, Weilong, et al.
Published: (2026)
Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks
by: Götz, Philipp, et al.
Published: (2025)
by: Götz, Philipp, et al.
Published: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
by: Xin, Detai, et al.
Published: (2024)
by: Xin, Detai, et al.
Published: (2024)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
by: Huang, Wen-Chin, et al.
Published: (2026)
by: Huang, Wen-Chin, et al.
Published: (2026)
Navigating PESQ: Up-to-Date Versions and Open Implementations
by: Torcoli, Matteo, et al.
Published: (2025)
by: Torcoli, Matteo, et al.
Published: (2025)
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
by: Shetu, Shrishti Saha, et al.
Published: (2025)
by: Shetu, Shrishti Saha, et al.
Published: (2025)
Speech Separation using Neural Audio Codecs with Embedding Loss
by: Yip, Jia Qi, et al.
Published: (2024)
by: Yip, Jia Qi, et al.
Published: (2024)
Neural Directional Filtering Using a Compact Microphone Array
by: Huang, Weilong, et al.
Published: (2025)
by: Huang, Weilong, et al.
Published: (2025)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
by: Li, Hanzhao, et al.
Published: (2024)
by: Li, Hanzhao, et al.
Published: (2024)
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
by: Zheng, Youqiang, et al.
Published: (2024)
by: Zheng, Youqiang, et al.
Published: (2024)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
by: Shi, Jiatong, et al.
Published: (2024)
by: Shi, Jiatong, et al.
Published: (2024)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
by: Yang, Leyan, et al.
Published: (2026)
by: Yang, Leyan, et al.
Published: (2026)
Probing the Robustness Properties of Neural Speech Codecs
by: Tseng, Wei-Cheng, et al.
Published: (2025)
by: Tseng, Wei-Cheng, et al.
Published: (2025)
SpatialCodec: Neural Spatial Speech Coding
by: Xu, Zhongweiyang, et al.
Published: (2023)
by: Xu, Zhongweiyang, et al.
Published: (2023)
A Neural Speech Codec for Noise Robust Speech Coding
by: Huang, Jiayi, et al.
Published: (2023)
by: Huang, Jiayi, et al.
Published: (2023)
Stereo Reproduction in the Presence of Sample Rate Offsets
by: Korse, Srikanth, et al.
Published: (2025)
by: Korse, Srikanth, et al.
Published: (2025)
On the Language and Gender Biases in PSTN, VoIP and Neural Audio Codecs
by: Altwlkany, Kemal, et al.
Published: (2025)
by: Altwlkany, Kemal, et al.
Published: (2025)
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages
by: Girish, et al.
Published: (2026)
by: Girish, et al.
Published: (2026)
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
by: Wang, Hankun, et al.
Published: (2025)
by: Wang, Hankun, et al.
Published: (2025)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
by: Chen, Xuanjun, et al.
Published: (2025)
by: Chen, Xuanjun, et al.
Published: (2025)
VoxATtack: A Multimodal Attack on Voice Anonymization Systems
by: Aloradi, Ahmad, et al.
Published: (2025)
by: Aloradi, Ahmad, et al.
Published: (2025)
Room Impulse Response Completion Using Signal-Prediction Diffusion Models Conditioned on Simulated Early Reflections
by: Xu, Zeyu, et al.
Published: (2026)
by: Xu, Zeyu, et al.
Published: (2026)
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
by: Korse, Srikanth, et al.
Published: (2025)
by: Korse, Srikanth, et al.
Published: (2025)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
by: Tseng, Wei-Cheng, et al.
Published: (2025)
by: Tseng, Wei-Cheng, et al.
Published: (2025)
Similar Items
-
Assessing the Impact of Noise and Speech Enhancement on the Intelligibility of Speech Codecs
by: Behringer, Lyonel, et al.
Published: (2026) -
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
by: Lakshminarayana, Kishor Kayyar, et al.
Published: (2025) -
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
by: Brendel, Andreas, et al.
Published: (2024) -
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
by: Halimeh, Mhd Modar, et al.
Published: (2025) -
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
by: Grundhuber, Philipp, et al.
Published: (2025)