Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Ye-Xin, Ai, Yang, Sheng, Zheng-Yan, Ling, Zhen-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
von: Liu, Fei, et al.
Veröffentlicht: (2024)
von: Liu, Fei, et al.
Veröffentlicht: (2024)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Universal Preference-Score-based Pairwise Speech Quality Assessment
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Vector Quantized Diffusion Model Based Speech Bandwidth Extension
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension
von: Gupta, Kishan, et al.
Veröffentlicht: (2025)
von: Gupta, Kishan, et al.
Veröffentlicht: (2025)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
Vision-Integrated High-Quality Neural Speech Coding
von: Guo, Yao, et al.
Veröffentlicht: (2025)
von: Guo, Yao, et al.
Veröffentlicht: (2025)
Fast and Flexible Audio Bandwidth Extension via Vocos
von: Sharma, Yatharth
Veröffentlicht: (2026)
von: Sharma, Yatharth
Veröffentlicht: (2026)
CIS-BWE: Chaos-Informed Speech Bandwidth Extension
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
von: Tamiti, Tarikul Islam, et al.
Veröffentlicht: (2025)
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
von: Niu, Rui, et al.
Veröffentlicht: (2025)
von: Niu, Rui, et al.
Veröffentlicht: (2025)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
MambaRate: Speech Quality Assessment Across Different Sampling Rates
von: Kakoulidis, Panos, et al.
Veröffentlicht: (2025)
von: Kakoulidis, Panos, et al.
Veröffentlicht: (2025)
Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
von: Bargum, Anders R., et al.
Veröffentlicht: (2024)
von: Bargum, Anders R., et al.
Veröffentlicht: (2024)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Configurable EBEN: Extreme Bandwidth Extension Network to enhance body-conducted speech capture
von: Hauret, Julien, et al.
Veröffentlicht: (2023)
von: Hauret, Julien, et al.
Veröffentlicht: (2023)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
Voice Attribute Editing with Text Prompt
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2024)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2024)
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models
von: Yang, Ningyuan, et al.
Veröffentlicht: (2026)
von: Yang, Ningyuan, et al.
Veröffentlicht: (2026)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning
von: Zhang, Daning, et al.
Veröffentlicht: (2025)
von: Zhang, Daning, et al.
Veröffentlicht: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
Rate-Aware Learned Speech Compression
von: Xu, Jun, et al.
Veröffentlicht: (2025)
von: Xu, Jun, et al.
Veröffentlicht: (2025)
Masked Audio Modeling with CLAP and Multi-Objective Learning
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024) -
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023) -
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
von: Ai, Yang, et al.
Veröffentlicht: (2024) -
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024) -
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
von: Liu, Fei, et al.
Veröffentlicht: (2024)