Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908720086122496 |
|---|---|
| author | Ellinas, Nikolaos Vioni, Alexandra Kakoulidis, Panos Vamvoukakis, Georgios Christidou, Myrsini Markopoulos, Konstantinos Oh, Junkwang Jho, Gunu Hwang, Inchul Chalamandaris, Aimilios Tsiakoulis, Pirros |
| author_facet | Ellinas, Nikolaos Vioni, Alexandra Kakoulidis, Panos Vamvoukakis, Georgios Christidou, Myrsini Markopoulos, Konstantinos Oh, Junkwang Jho, Gunu Hwang, Inchul Chalamandaris, Aimilios Tsiakoulis, Pirros |
| contents | This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-based vocoder without requiring any additional training or changes to the model. This is achieved by directly modifying the cepstrum feature space in order to shift the harmonic structure to the desired target. The spectrogram magnitude is computed via the pseudo-inverse mel transform, then converted to the cepstrum by applying DCT. In this domain, the cepstral peak is shifted without having to estimate its position and the modified mel is recomputed by applying IDCT and mel-filterbank. These pitch-shifted mel-spectrogram features can be converted to speech with any compatible vocoder. The proposed method is validated experimentally with objective and subjective metrics on various state-of-the-art neural vocoders as well as in comparison with traditional pitch modification methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_16519 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders Ellinas, Nikolaos Vioni, Alexandra Kakoulidis, Panos Vamvoukakis, Georgios Christidou, Myrsini Markopoulos, Konstantinos Oh, Junkwang Jho, Gunu Hwang, Inchul Chalamandaris, Aimilios Tsiakoulis, Pirros Sound Machine Learning Audio and Speech Processing This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-based vocoder without requiring any additional training or changes to the model. This is achieved by directly modifying the cepstrum feature space in order to shift the harmonic structure to the desired target. The spectrogram magnitude is computed via the pseudo-inverse mel transform, then converted to the cepstrum by applying DCT. In this domain, the cepstral peak is shifted without having to estimate its position and the modified mel is recomputed by applying IDCT and mel-filterbank. These pitch-shifted mel-spectrogram features can be converted to speech with any compatible vocoder. The proposed method is validated experimentally with objective and subjective metrics on various state-of-the-art neural vocoders as well as in comparison with traditional pitch modification methods. |
| title | Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders |
| topic | Sound Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2512.16519 |