Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914268548431872 |
|---|---|
| author | Al-Radhi, Mohammed Salah Larbi, Riad Bartalis, Mátyás Németh, Géza |
| author_facet | Al-Radhi, Mohammed Salah Larbi, Riad Bartalis, Mátyás Németh, Géza |
| contents | Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance voiced segment encoding and directly predicts complex spectral components for waveform synthesis via inverse STFT. Unlike mel-spectrogram-based approaches, our design jointly models magnitude and phase, ensuring phase coherence and improved pitch fidelity. To further align with perceptual quality, we adopt a multi-objective training strategy that integrates adversarial, spectral, and phase-aware losses. Experiments on benchmark datasets demonstrate consistent gains over HiFi-GAN and AutoVocoder: F0 RMSE reduced by 22 percent, voiced/unvoiced error lowered by 18 percent, and MOS scores improved by 0.15. These results show that prosody-guided attention combined with direct complex spectrum modeling yields more natural, pitch-accurate, and robust synthetic speech, setting a strong foundation for expressive neural vocoding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_14472 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum Al-Radhi, Mohammed Salah Larbi, Riad Bartalis, Mátyás Németh, Géza Sound Artificial Intelligence Computation and Language Neural vocoders are central to speech synthesis; despite their success, most still suffer from limited prosody modeling and inaccurate phase reconstruction. We propose a vocoder that introduces prosody-guided harmonic attention to enhance voiced segment encoding and directly predicts complex spectral components for waveform synthesis via inverse STFT. Unlike mel-spectrogram-based approaches, our design jointly models magnitude and phase, ensuring phase coherence and improved pitch fidelity. To further align with perceptual quality, we adopt a multi-objective training strategy that integrates adversarial, spectral, and phase-aware losses. Experiments on benchmark datasets demonstrate consistent gains over HiFi-GAN and AutoVocoder: F0 RMSE reduced by 22 percent, voiced/unvoiced error lowered by 18 percent, and MOS scores improved by 0.15. These results show that prosody-guided attention combined with direct complex spectrum modeling yields more natural, pitch-accurate, and robust synthetic speech, setting a strong foundation for expressive neural vocoding. |
| title | Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum |
| topic | Sound Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2601.14472 |