On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Halimeh, Mhd Modar, Torcoli, Matteo, Grundhuber, Philipp, Habets, Emanuël A. P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
von: Halimeh, Mhd Modar, et al.
Veröffentlicht: (2024)
von: Halimeh, Mhd Modar, et al.
Veröffentlicht: (2024)
Navigating PESQ: Up-to-Date Versions and Open Implementations
von: Torcoli, Matteo, et al.
Veröffentlicht: (2025)
von: Torcoli, Matteo, et al.
Veröffentlicht: (2025)
Robust Speech Activity Detection in the Presence of Singing Voice
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
ODAQ: Open Dataset of Audio Quality
von: Torcoli, Matteo, et al.
Veröffentlicht: (2023)
von: Torcoli, Matteo, et al.
Veröffentlicht: (2023)
Speech Loudness in Broadcasting and Streaming
von: Torcoli, Matteo, et al.
Veröffentlicht: (2024)
von: Torcoli, Matteo, et al.
Veröffentlicht: (2024)
Neural Directional Filtering Using a Compact Microphone Array
von: Huang, Weilong, et al.
Veröffentlicht: (2025)
von: Huang, Weilong, et al.
Veröffentlicht: (2025)
Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array
von: Wechsler, Julian, et al.
Veröffentlicht: (2024)
von: Wechsler, Julian, et al.
Veröffentlicht: (2024)
Benchmarking Neural Speech Codec Intelligibility with SITool
von: Leschanowsky, Anna, et al.
Veröffentlicht: (2025)
von: Leschanowsky, Anna, et al.
Veröffentlicht: (2025)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
Expanding and Analyzing ODAQ -- the Open Dataset of Audio Quality
von: Dick, Sascha, et al.
Veröffentlicht: (2025)
von: Dick, Sascha, et al.
Veröffentlicht: (2025)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Dynamic Slimmable Networks for Efficient Speech Separation
von: Elminshawi, Mohamed, et al.
Veröffentlicht: (2025)
von: Elminshawi, Mohamed, et al.
Veröffentlicht: (2025)
Blind Acoustic Parameter Estimation Through Task-Agnostic Embeddings Using Latent Approximations
von: Götz, Philipp, et al.
Veröffentlicht: (2024)
von: Götz, Philipp, et al.
Veröffentlicht: (2024)
Neural Directional Filtering with Configurable Directivity Pattern at Inference
von: Huang, Weilong, et al.
Veröffentlicht: (2025)
von: Huang, Weilong, et al.
Veröffentlicht: (2025)
Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks
von: Götz, Philipp, et al.
Veröffentlicht: (2025)
von: Götz, Philipp, et al.
Veröffentlicht: (2025)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
von: Lakshminarayana, Kishor Kayyar, et al.
Veröffentlicht: (2025)
von: Lakshminarayana, Kishor Kayyar, et al.
Veröffentlicht: (2025)
Sample Rate Offset Compensated Acoustic Echo Cancellation For Multi-Device Scenarios
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction
von: Huang, Weilong, et al.
Veröffentlicht: (2026)
von: Huang, Weilong, et al.
Veröffentlicht: (2026)
Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
von: Zhang, En-Wei, et al.
Veröffentlicht: (2025)
von: Zhang, En-Wei, et al.
Veröffentlicht: (2025)
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
Personalized Neural Speech Codec
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Investigating the impact of stereo processing -- a study for extending the Open Dataset of Audio Quality (ODAQ)
von: Dick, Sascha, et al.
Veröffentlicht: (2025)
von: Dick, Sascha, et al.
Veröffentlicht: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
Stereo Reproduction in the Presence of Sample Rate Offsets
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
von: Korse, Srikanth, et al.
Veröffentlicht: (2025)
Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
RepCodec: A Speech Representation Codec for Speech Tokenization
von: Huang, Zhichao, et al.
Veröffentlicht: (2023)
von: Huang, Zhichao, et al.
Veröffentlicht: (2023)
Speech Separation using Neural Audio Codecs with Embedding Loss
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
von: Li, Hanzhao, et al.
Veröffentlicht: (2024)
von: Li, Hanzhao, et al.
Veröffentlicht: (2024)
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
von: Yang, Leyan, et al.
Veröffentlicht: (2026)
von: Yang, Leyan, et al.
Veröffentlicht: (2026)
Probing the Robustness Properties of Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
SpatialCodec: Neural Spatial Speech Coding
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
A Neural Speech Codec for Noise Robust Speech Coding
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025) -
ConcateNet: Dialogue Separation Using Local And Global Feature Concatenation
von: Halimeh, Mhd Modar, et al.
Veröffentlicht: (2024) -
Navigating PESQ: Up-to-Date Versions and Open Implementations
von: Torcoli, Matteo, et al.
Veröffentlicht: (2025) -
Robust Speech Activity Detection in the Presence of Singing Voice
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025) -
ODAQ: Open Dataset of Audio Quality
von: Torcoli, Matteo, et al.
Veröffentlicht: (2023)