Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kolani, Yakov, Melichov, Maxim, Calev, Cobi, Alper, Morris
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911200693977088
author Kolani, Yakov
Melichov, Maxim
Calev, Cobi
Alper, Morris
author_facet Kolani, Yakov
Melichov, Maxim
Calev, Cobi
Alper, Morris
contents Real-time text-to-speech (TTS) for Modern Hebrew is challenging due to the language's orthographic complexity. Existing solutions ignore crucial phonetic features such as stress that remain underspecified even when vowel marks are added. To address these limitations, we introduce Phonikud, a lightweight, open-source Hebrew grapheme-to-phoneme (G2P) system that outputs fully-specified IPA transcriptions. Our approach adapts an existing diacritization model with lightweight adaptors, incurring negligible additional latency. We also contribute the ILSpeech dataset of transcribed Hebrew speech with IPA annotations, serving as a benchmark for Hebrew G2P, as training data for TTS systems, and enabling audio-to-IPA for evaluating TTS performance while capturing important phonetic details. Our results demonstrate that Phonikud G2P conversion more accurately predicts phonemes from Hebrew text compared to prior methods, and that this enables training of effective real-time Hebrew TTS models with superior speed-accuracy trade-offs. We release our code, data, and models at https: //phonikud.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12311
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech
Kolani, Yakov
Melichov, Maxim
Calev, Cobi
Alper, Morris
Computation and Language
Sound
Audio and Speech Processing
Real-time text-to-speech (TTS) for Modern Hebrew is challenging due to the language's orthographic complexity. Existing solutions ignore crucial phonetic features such as stress that remain underspecified even when vowel marks are added. To address these limitations, we introduce Phonikud, a lightweight, open-source Hebrew grapheme-to-phoneme (G2P) system that outputs fully-specified IPA transcriptions. Our approach adapts an existing diacritization model with lightweight adaptors, incurring negligible additional latency. We also contribute the ILSpeech dataset of transcribed Hebrew speech with IPA annotations, serving as a benchmark for Hebrew G2P, as training data for TTS systems, and enabling audio-to-IPA for evaluating TTS performance while capturing important phonetic details. Our results demonstrate that Phonikud G2P conversion more accurately predicts phonemes from Hebrew text compared to prior methods, and that this enables training of effective real-time Hebrew TTS models with superior speed-accuracy trade-offs. We release our code, data, and models at https: //phonikud.github.io.
title Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.12311