Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fetrat, Mahta, Navabi, Donya, Dehghanian, Zahra, Abolghasemi, Morteza, Rabiee, Hamid R.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915671349133312
author Fetrat, Mahta
Navabi, Donya
Dehghanian, Zahra
Abolghasemi, Morteza
Rabiee, Hamid R.
author_facet Fetrat, Mahta
Navabi, Donya
Dehghanian, Zahra
Abolghasemi, Morteza
Rabiee, Hamid R.
contents Lightweight, real-time text-to-speech systems are crucial for accessibility. However, the most efficient TTS models often rely on lightweight phonemizers that struggle with context-dependent challenges. In contrast, more advanced phonemizers with a deeper linguistic understanding typically incur high computational costs, which prevents real-time performance. This paper examines the trade-off between phonemization quality and inference speed in G2P-aided TTS systems, introducing a practical framework to bridge this gap. We propose lightweight strategies for context-aware phonemization and a service-oriented TTS architecture that executes these modules as independent services. This design decouples heavy context-aware components from the core TTS engine, effectively breaking the latency barrier and enabling real-time use of high-quality phonemization models. Experimental results confirm that the proposed system improves pronunciation soundness and linguistic accuracy while maintaining real-time responsiveness, making it well-suited for offline and end-device TTS applications.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08006
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
Fetrat, Mahta
Navabi, Donya
Dehghanian, Zahra
Abolghasemi, Morteza
Rabiee, Hamid R.
Sound
Computation and Language
Audio and Speech Processing
Lightweight, real-time text-to-speech systems are crucial for accessibility. However, the most efficient TTS models often rely on lightweight phonemizers that struggle with context-dependent challenges. In contrast, more advanced phonemizers with a deeper linguistic understanding typically incur high computational costs, which prevents real-time performance. This paper examines the trade-off between phonemization quality and inference speed in G2P-aided TTS systems, introducing a practical framework to bridge this gap. We propose lightweight strategies for context-aware phonemization and a service-oriented TTS architecture that executes these modules as independent services. This design decouples heavy context-aware components from the core TTS engine, effectively breaking the latency barrier and enabling real-time use of high-quality phonemization models. Experimental results confirm that the proposed system improves pronunciation soundness and linguistic accuracy while maintaining real-time responsiveness, making it well-suited for offline and end-device TTS applications.
title Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2512.08006