Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Rui, Lin, Xiaolong, Liu, Jiawang, Huang, Shixi, Zhan, Zhenpeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913885284466688
author Hu, Rui
Lin, Xiaolong
Liu, Jiawang
Huang, Shixi
Zhan, Zhenpeng
author_facet Hu, Rui
Lin, Xiaolong
Liu, Jiawang
Huang, Shixi
Zhan, Zhenpeng
contents In this paper, we propose a method for annotating phonemic and prosodic labels on a given audio-transcript pair, aimed at constructing Japanese text-to-speech (TTS) datasets. Our approach involves fine-tuning a large-scale pre-trained automatic speech recognition (ASR) model, conditioned on ground truth transcripts, to simultaneously output phrase-level graphemes and annotation labels. To further correct errors in phonemic labeling, we employ a decoding strategy that utilizes dictionary prior knowledge. The objective evaluation results demonstrate that our proposed method outperforms previous approaches relying solely on text or audio. The subjective evaluation results indicate that the naturalness of speech synthesized by the TTS model, trained with labels annotated using our method, is comparable to that of a model trained with manual annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07646
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
Hu, Rui
Lin, Xiaolong
Liu, Jiawang
Huang, Shixi
Zhan, Zhenpeng
Computation and Language
Sound
Audio and Speech Processing
In this paper, we propose a method for annotating phonemic and prosodic labels on a given audio-transcript pair, aimed at constructing Japanese text-to-speech (TTS) datasets. Our approach involves fine-tuning a large-scale pre-trained automatic speech recognition (ASR) model, conditioned on ground truth transcripts, to simultaneously output phrase-level graphemes and annotation labels. To further correct errors in phonemic labeling, we employ a decoding strategy that utilizes dictionary prior knowledge. The objective evaluation results demonstrate that our proposed method outperforms previous approaches relying solely on text or audio. The subjective evaluation results indicate that the naturalness of speech synthesized by the TTS model, trained with labels annotated using our method, is comparable to that of a model trained with manual annotations.
title Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.07646