From Hallucination to Articulation: Language Model-Driven Losses for Ultra Low-Bitrate Neural Speech Coding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yi, Jayeon, Kim, Minje |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding
von: Zhao, Mingyu, et al.
Veröffentlicht: (2026)
von: Zhao, Mingyu, et al.
Veröffentlicht: (2026)
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
von: Ren, Yanzhou, et al.
Veröffentlicht: (2026)
von: Ren, Yanzhou, et al.
Veröffentlicht: (2026)
Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2024)
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2024)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
von: Kim, Minje, et al.
Veröffentlicht: (2024)
von: Kim, Minje, et al.
Veröffentlicht: (2024)
CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
von: Yang, Leyan, et al.
Veröffentlicht: (2026)
von: Yang, Leyan, et al.
Veröffentlicht: (2026)
Hyperbolic Distance-Based Speech Separation
von: Petermann, Darius, et al.
Veröffentlicht: (2024)
von: Petermann, Darius, et al.
Veröffentlicht: (2024)
Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
von: Shechtman, Slava, et al.
Veröffentlicht: (2024)
von: Shechtman, Slava, et al.
Veröffentlicht: (2024)
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
von: Brendel, Andreas, et al.
Veröffentlicht: (2024)
von: Brendel, Andreas, et al.
Veröffentlicht: (2024)
Personalized Neural Speech Codec
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding
von: Yang, Peiji, et al.
Veröffentlicht: (2024)
von: Yang, Peiji, et al.
Veröffentlicht: (2024)
MuCodec: Ultra Low-Bitrate Music Codec
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
DDD: A Perceptually Superior Low-Response-Time DNN-based Declipper
von: Yi, Jayeon, et al.
Veröffentlicht: (2024)
von: Yi, Jayeon, et al.
Veröffentlicht: (2024)
CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2026)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2026)
Perceptual Audio Coding: A 40-Year Historical Perspective
von: Herre, Jürgen, et al.
Veröffentlicht: (2025)
von: Herre, Jürgen, et al.
Veröffentlicht: (2025)
TGIF: Talker Group-Informed Familiarization of Target Speaker Extraction
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2025)
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2025)
Adaptive Deterministic Flow Matching for Target Speaker Extraction
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2025)
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2025)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement
von: Li, Jingyu, et al.
Veröffentlicht: (2025)
von: Li, Jingyu, et al.
Veröffentlicht: (2025)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
A Comparative Analysis of Poetry Reading Audio: Singing, Narrating, or Somewhere In Between?
von: Choi, Kahyun, et al.
Veröffentlicht: (2024)
von: Choi, Kahyun, et al.
Veröffentlicht: (2024)
Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
von: Close, George, et al.
Veröffentlicht: (2024)
von: Close, George, et al.
Veröffentlicht: (2024)
StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
von: Della Libera, Luca, et al.
Veröffentlicht: (2025)
Combolutional Neural Networks
von: Churchwell, Cameron, et al.
Veröffentlicht: (2025)
von: Churchwell, Cameron, et al.
Veröffentlicht: (2025)
PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
Speech Separation using Neural Audio Codecs with Embedding Loss
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
Gencho: Room Impulse Response Generation from Reverberant Speech and Text via Diffusion Transformers
von: Lin, Jackie, et al.
Veröffentlicht: (2026)
von: Lin, Jackie, et al.
Veröffentlicht: (2026)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
von: Kuznetsova, Anastasia, et al.
Veröffentlicht: (2025)
von: Kuznetsova, Anastasia, et al.
Veröffentlicht: (2025)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding
von: Zhao, Mingyu, et al.
Veröffentlicht: (2026) -
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
von: Ren, Yanzhou, et al.
Veröffentlicht: (2026) -
Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026) -
Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2024) -
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
von: Kim, Minje, et al.
Veröffentlicht: (2024)