Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ohnaka, Hien, Shirahata, Yuma, Kawamura, Masaya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
von: Ohnaka, Hien, et al.
Veröffentlicht: (2025)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2025)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
Universal Score-based Speech Enhancement with High Content Preservation
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
von: Kawamura, Masaya, et al.
Veröffentlicht: (2024)
von: Kawamura, Masaya, et al.
Veröffentlicht: (2024)
Unsupervised speech enhancement with spectral kurtosis and double deep priors
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
Neural Vocoders as Speech Enhancers
von: Li, Andong, et al.
Veröffentlicht: (2025)
von: Li, Andong, et al.
Veröffentlicht: (2025)
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
von: Terashima, Ryo, et al.
Veröffentlicht: (2025)
von: Terashima, Ryo, et al.
Veröffentlicht: (2025)
Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis
von: Agrawal, Prabhav, et al.
Veröffentlicht: (2024)
von: Agrawal, Prabhav, et al.
Veröffentlicht: (2024)
Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Vision-Integrated High-Quality Neural Speech Coding
von: Guo, Yao, et al.
Veröffentlicht: (2025)
von: Guo, Yao, et al.
Veröffentlicht: (2025)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2024)
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2024)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems
von: Liu, Jeongmin, et al.
Veröffentlicht: (2024)
von: Liu, Jeongmin, et al.
Veröffentlicht: (2024)
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
Mitigating Language Mismatch in SSL-Based Speaker Anonymization
von: Zhang, Zhe, et al.
Veröffentlicht: (2025)
von: Zhang, Zhe, et al.
Veröffentlicht: (2025)
PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
von: Ohnaka, Hien, et al.
Veröffentlicht: (2025) -
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025) -
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024) -
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025) -
Universal Score-based Speech Enhancement with High Content Preservation
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)