JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kondo, Yuto, Kameoka, Hirokazu, Tanaka, Kou, Kaneko, Takuhiro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911058563694592
author Kondo, Yuto
Kameoka, Hirokazu
Tanaka, Kou
Kaneko, Takuhiro
author_facet Kondo, Yuto
Kameoka, Hirokazu
Tanaka, Kou
Kaneko, Takuhiro
contents We construct Japanese Idol Speech Corpus (JIS) to advance research in speech generation AI, including text-to-speech synthesis (TTS) and voice conversion (VC). JIS will facilitate more rigorous evaluations of speaker similarity in TTS and VC systems since all speakers in JIS belong to a highly specific category: "young female live idols" in Japan, and each speaker is identified by a stage name, enabling researchers to recruit listeners familiar with these idols for listening experiments. With its unique speaker attributes, JIS will foster compelling research, including generating voices tailored to listener preferences-an area not yet widely studied. JIS will be distributed free of charge to promote research in speech generation AI, with usage restricted to non-commercial, basic research. We describe the construction of JIS, provide an overview of Japanese live idol culture to support effective and ethical use of JIS, and offer a basic analysis to guide application of JIS.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18296
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
Kondo, Yuto
Kameoka, Hirokazu
Tanaka, Kou
Kaneko, Takuhiro
Sound
Audio and Speech Processing
We construct Japanese Idol Speech Corpus (JIS) to advance research in speech generation AI, including text-to-speech synthesis (TTS) and voice conversion (VC). JIS will facilitate more rigorous evaluations of speaker similarity in TTS and VC systems since all speakers in JIS belong to a highly specific category: "young female live idols" in Japan, and each speaker is identified by a stage name, enabling researchers to recruit listeners familiar with these idols for listening experiments. With its unique speaker attributes, JIS will foster compelling research, including generating voices tailored to listener preferences-an area not yet widely studied. JIS will be distributed free of charge to promote research in speech generation AI, with usage restricted to non-commercial, basic research. We describe the construction of JIS, provide an overview of Japanese live idol culture to support effective and ethical use of JIS, and offer a basic analysis to guide application of JIS.
title JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.18296