A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, En-Wei, Du, Hui-Peng, Jiang, Xiao-Hang, Ai, Yang, Ling, Zhen-Hua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement
por: Jiang, Xiao-Hang, et al.
Publicado: (2026)
por: Jiang, Xiao-Hang, et al.
Publicado: (2026)
CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction
por: Du, Hui-Peng, et al.
Publicado: (2026)
por: Du, Hui-Peng, et al.
Publicado: (2026)
Vision-Integrated High-Quality Neural Speech Coding
por: Guo, Yao, et al.
Publicado: (2025)
por: Guo, Yao, et al.
Publicado: (2025)
ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs
por: Zheng, Rui-Chen, et al.
Publicado: (2024)
por: Zheng, Rui-Chen, et al.
Publicado: (2024)
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
por: Du, Hui-Peng, et al.
Publicado: (2025)
por: Du, Hui-Peng, et al.
Publicado: (2025)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction
por: Du, Hui-Peng, et al.
Publicado: (2026)
por: Du, Hui-Peng, et al.
Publicado: (2026)
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
por: Ai, Yang, et al.
Publicado: (2024)
por: Ai, Yang, et al.
Publicado: (2024)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
por: Du, Hui-Peng, et al.
Publicado: (2024)
por: Du, Hui-Peng, et al.
Publicado: (2024)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
por: Lu, Ye-Xin, et al.
Publicado: (2024)
por: Lu, Ye-Xin, et al.
Publicado: (2024)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
por: Ai, Yang, et al.
Publicado: (2024)
por: Ai, Yang, et al.
Publicado: (2024)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation
por: Zheng, Rui-Chen, et al.
Publicado: (2025)
por: Zheng, Rui-Chen, et al.
Publicado: (2025)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
por: Lu, Ye-Xin, et al.
Publicado: (2023)
por: Lu, Ye-Xin, et al.
Publicado: (2023)
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2024)
por: Lu, Ye-Xin, et al.
Publicado: (2024)
MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement
por: Li, Jingyu, et al.
Publicado: (2025)
por: Li, Jingyu, et al.
Publicado: (2025)
Universal Preference-Score-based Pairwise Speech Quality Assessment
por: Shi, Yu-Fei, et al.
Publicado: (2025)
por: Shi, Yu-Fei, et al.
Publicado: (2025)
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
por: Ren, Yanzhou, et al.
Publicado: (2026)
por: Ren, Yanzhou, et al.
Publicado: (2026)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
por: Du, Hui-Peng, et al.
Publicado: (2024)
por: Du, Hui-Peng, et al.
Publicado: (2024)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
por: Du, Hui-Peng, et al.
Publicado: (2024)
por: Du, Hui-Peng, et al.
Publicado: (2024)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
por: Xin, Detai, et al.
Publicado: (2024)
por: Xin, Detai, et al.
Publicado: (2024)
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
por: Du, Hui-Peng, et al.
Publicado: (2025)
por: Du, Hui-Peng, et al.
Publicado: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
por: Yang, Leyan, et al.
Publicado: (2026)
por: Yang, Leyan, et al.
Publicado: (2026)
A Low-Complexity Speech Codec Using Parametric Dithering for ASR
por: Murray, Ellison, et al.
Publicado: (2025)
por: Murray, Ellison, et al.
Publicado: (2025)
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
por: Halimeh, Mhd Modar, et al.
Publicado: (2025)
por: Halimeh, Mhd Modar, et al.
Publicado: (2025)
Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
por: Shi, Mohan, et al.
Publicado: (2025)
por: Shi, Mohan, et al.
Publicado: (2025)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
por: Shi, Yu-Fei, et al.
Publicado: (2024)
por: Shi, Yu-Fei, et al.
Publicado: (2024)
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
por: Zheng, Youqiang, et al.
Publicado: (2024)
por: Zheng, Youqiang, et al.
Publicado: (2024)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
por: Li, Hanzhao, et al.
Publicado: (2024)
por: Li, Hanzhao, et al.
Publicado: (2024)
MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra
por: Lu, Ye-Xin, et al.
Publicado: (2023)
por: Lu, Ye-Xin, et al.
Publicado: (2023)
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
por: Liu, Fei, et al.
Publicado: (2024)
por: Liu, Fei, et al.
Publicado: (2024)
Probing the Robustness Properties of Neural Speech Codecs
por: Tseng, Wei-Cheng, et al.
Publicado: (2025)
por: Tseng, Wei-Cheng, et al.
Publicado: (2025)
SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding
por: Zhao, Mingyu, et al.
Publicado: (2026)
por: Zhao, Mingyu, et al.
Publicado: (2026)
A Neural Speech Codec for Noise Robust Speech Coding
por: Huang, Jiayi, et al.
Publicado: (2023)
por: Huang, Jiayi, et al.
Publicado: (2023)
PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios
por: Wan, Zixiang, et al.
Publicado: (2025)
por: Wan, Zixiang, et al.
Publicado: (2025)
Benchmarking Neural Speech Codec Intelligibility with SITool
por: Leschanowsky, Anna, et al.
Publicado: (2025)
por: Leschanowsky, Anna, et al.
Publicado: (2025)
SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features
por: Shi, Yu-Fei, et al.
Publicado: (2024)
por: Shi, Yu-Fei, et al.
Publicado: (2024)
Personalized Neural Speech Codec
por: Jang, Inseon, et al.
Publicado: (2024)
por: Jang, Inseon, et al.
Publicado: (2024)
Ejemplares similares
-
CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement
por: Jiang, Xiao-Hang, et al.
Publicado: (2026) -
CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction
por: Du, Hui-Peng, et al.
Publicado: (2026) -
Vision-Integrated High-Quality Neural Speech Coding
por: Guo, Yao, et al.
Publicado: (2025) -
ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs
por: Zheng, Rui-Chen, et al.
Publicado: (2024) -
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
por: Du, Hui-Peng, et al.
Publicado: (2025)