Scalable Neural Vocoder from Range-Null Space Decomposition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Andong, Lei, Tong, Sun, Zhihang, Chen, Rilin, Li, Xiaodong, Yu, Dong, Zheng, Chengshi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Neural Vocoder from Range-Null Space Decomposition
von: Li, Andong, et al.
Veröffentlicht: (2025)
von: Li, Andong, et al.
Veröffentlicht: (2025)
BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
von: Li, Andong, et al.
Veröffentlicht: (2025)
von: Li, Andong, et al.
Veröffentlicht: (2025)
Neural Vocoders as Speech Enhancers
von: Li, Andong, et al.
Veröffentlicht: (2025)
von: Li, Andong, et al.
Veröffentlicht: (2025)
SMRU: Split-and-Merge Recurrent-based UNet for Acoustic Echo Cancellation and Noise Suppression
von: Sun, Zhihang, et al.
Veröffentlicht: (2024)
von: Sun, Zhihang, et al.
Veröffentlicht: (2024)
Target matching based generative model for speech enhancement
von: Wang, Taihui, et al.
Veröffentlicht: (2025)
von: Wang, Taihui, et al.
Veröffentlicht: (2025)
GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks
von: Dai, Lingling, et al.
Veröffentlicht: (2026)
von: Dai, Lingling, et al.
Veröffentlicht: (2026)
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing
von: Liang, Yifan, et al.
Veröffentlicht: (2025)
von: Liang, Yifan, et al.
Veröffentlicht: (2025)
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2025)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2025)
Gen-SER: When the generative model meets speech emotion recognition
von: Wang, Taihui, et al.
Veröffentlicht: (2026)
von: Wang, Taihui, et al.
Veröffentlicht: (2026)
BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement
von: Fan, Cunhang, et al.
Veröffentlicht: (2024)
von: Fan, Cunhang, et al.
Veröffentlicht: (2024)
Video-to-Audio Generation with Fine-grained Temporal Semantics
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
High-Fidelity Music Vocoder using Neural Audio Codecs
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
End-to-end multi-channel speaker extraction and binaural speech synthesis
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
von: Xie, Zeyu, et al.
Veröffentlicht: (2026)
von: Xie, Zeyu, et al.
Veröffentlicht: (2026)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
von: Ren, Yong, et al.
Veröffentlicht: (2024)
von: Ren, Yong, et al.
Veröffentlicht: (2024)
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
Global Rotation Equivariant Phase Modeling for Speech Enhancement with Deep Magnitude-Phase Interaction
von: Wang, Chengzhong, et al.
Veröffentlicht: (2026)
von: Wang, Chengzhong, et al.
Veröffentlicht: (2026)
Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2026)
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2026)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis
von: Agrawal, Prabhav, et al.
Veröffentlicht: (2024)
von: Agrawal, Prabhav, et al.
Veröffentlicht: (2024)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
LCM-SVC: Latent Diffusion Model Based Singing Voice Conversion with Inference Acceleration via Latent Consistency Distillation
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
Vocoder-Projected Feature Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
von: Dai, Lingling, et al.
Veröffentlicht: (2025)
von: Dai, Lingling, et al.
Veröffentlicht: (2025)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
von: Ohnaka, Hien, et al.
Veröffentlicht: (2026)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2026)
LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
von: Yang, Kang, et al.
Veröffentlicht: (2025)
von: Yang, Kang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning Neural Vocoder from Range-Null Space Decomposition
von: Li, Andong, et al.
Veröffentlicht: (2025) -
BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
von: Li, Andong, et al.
Veröffentlicht: (2025) -
Neural Vocoders as Speech Enhancers
von: Li, Andong, et al.
Veröffentlicht: (2025) -
SMRU: Split-and-Merge Recurrent-based UNet for Acoustic Echo Cancellation and Noise Suppression
von: Sun, Zhihang, et al.
Veröffentlicht: (2024) -
Target matching based generative model for speech enhancement
von: Wang, Taihui, et al.
Veröffentlicht: (2025)