Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction
Fuente:
arXiv
Salvato in:
| Autori principali: | Du, Hui-Peng, Ai, Yang, Jiang, Xiao-Hang, Tian, Yuan, Ling, Zhen-Hua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
di: Du, Hui-Peng, et al.
Pubblicazione: (2025)
di: Du, Hui-Peng, et al.
Pubblicazione: (2025)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
di: Jiang, Xiao-Hang, et al.
Pubblicazione: (2024)
di: Jiang, Xiao-Hang, et al.
Pubblicazione: (2024)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
di: Du, Hui-Peng, et al.
Pubblicazione: (2024)
di: Du, Hui-Peng, et al.
Pubblicazione: (2024)
CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement
di: Jiang, Xiao-Hang, et al.
Pubblicazione: (2026)
di: Jiang, Xiao-Hang, et al.
Pubblicazione: (2026)
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
di: Du, Hui-Peng, et al.
Pubblicazione: (2025)
di: Du, Hui-Peng, et al.
Pubblicazione: (2025)
CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction
di: Du, Hui-Peng, et al.
Pubblicazione: (2026)
di: Du, Hui-Peng, et al.
Pubblicazione: (2026)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
di: Du, Hui-Peng, et al.
Pubblicazione: (2024)
di: Du, Hui-Peng, et al.
Pubblicazione: (2024)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
di: Jiang, Xiao-Hang, et al.
Pubblicazione: (2024)
di: Jiang, Xiao-Hang, et al.
Pubblicazione: (2024)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
di: Ai, Yang, et al.
Pubblicazione: (2024)
di: Ai, Yang, et al.
Pubblicazione: (2024)
A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
di: Zhang, En-Wei, et al.
Pubblicazione: (2025)
di: Zhang, En-Wei, et al.
Pubblicazione: (2025)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
di: Guo, Hongming, et al.
Pubblicazione: (2024)
di: Guo, Hongming, et al.
Pubblicazione: (2024)
Vision-Integrated High-Quality Neural Speech Coding
di: Guo, Yao, et al.
Pubblicazione: (2025)
di: Guo, Yao, et al.
Pubblicazione: (2025)
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
di: Zhou, Rui, et al.
Pubblicazione: (2024)
di: Zhou, Rui, et al.
Pubblicazione: (2024)
Real-Time Streaming Mel Vocoding with Generative Flow Matching
di: Welker, Simon, et al.
Pubblicazione: (2025)
di: Welker, Simon, et al.
Pubblicazione: (2025)
SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding
di: Zhao, Mingyu, et al.
Pubblicazione: (2026)
di: Zhao, Mingyu, et al.
Pubblicazione: (2026)
From Hallucination to Articulation: Language Model-Driven Losses for Ultra Low-Bitrate Neural Speech Coding
di: Yi, Jayeon, et al.
Pubblicazione: (2026)
di: Yi, Jayeon, et al.
Pubblicazione: (2026)
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
di: Shao, Nian, et al.
Pubblicazione: (2025)
di: Shao, Nian, et al.
Pubblicazione: (2025)
DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
di: Ellinas, Nikolaos, et al.
Pubblicazione: (2025)
di: Ellinas, Nikolaos, et al.
Pubblicazione: (2025)
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
di: Ren, Yanzhou, et al.
Pubblicazione: (2026)
di: Ren, Yanzhou, et al.
Pubblicazione: (2026)
Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
di: Shechtman, Slava, et al.
Pubblicazione: (2024)
di: Shechtman, Slava, et al.
Pubblicazione: (2024)
Neural Vocoders as Speech Enhancers
di: Li, Andong, et al.
Pubblicazione: (2025)
di: Li, Andong, et al.
Pubblicazione: (2025)
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2026)
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2026)
Universal Preference-Score-based Pairwise Speech Quality Assessment
di: Shi, Yu-Fei, et al.
Pubblicazione: (2025)
di: Shi, Yu-Fei, et al.
Pubblicazione: (2025)
A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
di: Hu, Guoqiang, et al.
Pubblicazione: (2024)
di: Hu, Guoqiang, et al.
Pubblicazione: (2024)
Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
di: Guimarães, Heitor R., et al.
Pubblicazione: (2025)
di: Guimarães, Heitor R., et al.
Pubblicazione: (2025)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
di: Du, Hui-Peng, et al.
Pubblicazione: (2024)
di: Du, Hui-Peng, et al.
Pubblicazione: (2024)
ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs
di: Zheng, Rui-Chen, et al.
Pubblicazione: (2024)
di: Zheng, Rui-Chen, et al.
Pubblicazione: (2024)
FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
di: Lv, Yuanjun, et al.
Pubblicazione: (2024)
di: Lv, Yuanjun, et al.
Pubblicazione: (2024)
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2025)
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2025)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
di: Chae, Yunkee, et al.
Pubblicazione: (2025)
di: Chae, Yunkee, et al.
Pubblicazione: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
di: Xin, Detai, et al.
Pubblicazione: (2024)
di: Xin, Detai, et al.
Pubblicazione: (2024)
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
di: Ai, Yang, et al.
Pubblicazione: (2024)
di: Ai, Yang, et al.
Pubblicazione: (2024)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
di: Yang, Leyan, et al.
Pubblicazione: (2026)
di: Yang, Leyan, et al.
Pubblicazione: (2026)
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
di: Brendel, Andreas, et al.
Pubblicazione: (2024)
di: Brendel, Andreas, et al.
Pubblicazione: (2024)
MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra
di: Lu, Ye-Xin, et al.
Pubblicazione: (2023)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2023)
Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis
di: Agrawal, Prabhav, et al.
Pubblicazione: (2024)
di: Agrawal, Prabhav, et al.
Pubblicazione: (2024)
Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding
di: Yang, Peiji, et al.
Pubblicazione: (2024)
di: Yang, Peiji, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
di: Du, Hui-Peng, et al.
Pubblicazione: (2025) -
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
di: Jiang, Xiao-Hang, et al.
Pubblicazione: (2024) -
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
di: Du, Hui-Peng, et al.
Pubblicazione: (2024) -
CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement
di: Jiang, Xiao-Hang, et al.
Pubblicazione: (2026) -
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
di: Du, Hui-Peng, et al.
Pubblicazione: (2025)