Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
Fuente:
arXiv
Salvato in:
| Autori principali: | Shirahata, Yuma, Park, Byeongseon, Yamamoto, Ryuichi, Tachibana, Kentaro |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
di: Ohnaka, Hien, et al.
Pubblicazione: (2025)
di: Ohnaka, Hien, et al.
Pubblicazione: (2025)
CC-G2PnP: Streaming Grapheme-to-Phoneme and prosody with Conformer-CTC for unsegmented languages
di: Shirahata, Yuma, et al.
Pubblicazione: (2026)
di: Shirahata, Yuma, et al.
Pubblicazione: (2026)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
di: Yamamoto, Ryuichi, et al.
Pubblicazione: (2024)
di: Yamamoto, Ryuichi, et al.
Pubblicazione: (2024)
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
di: Kawamura, Masaya, et al.
Pubblicazione: (2024)
di: Kawamura, Masaya, et al.
Pubblicazione: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
di: Ma, Te, et al.
Pubblicazione: (2025)
di: Ma, Te, et al.
Pubblicazione: (2025)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
di: Garg, Abhinav, et al.
Pubblicazione: (2024)
di: Garg, Abhinav, et al.
Pubblicazione: (2024)
Disentangling segmental and prosodic factors to non-native speech comprehensibility
di: Quamer, Waris, et al.
Pubblicazione: (2024)
di: Quamer, Waris, et al.
Pubblicazione: (2024)
BabAR: from phoneme recognition to developmental measures of young children's speech production
di: Lavechin, Marvin, et al.
Pubblicazione: (2026)
di: Lavechin, Marvin, et al.
Pubblicazione: (2026)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
di: Kawamura, Masaya, et al.
Pubblicazione: (2025)
di: Kawamura, Masaya, et al.
Pubblicazione: (2025)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
di: Lyth, Dan, et al.
Pubblicazione: (2024)
di: Lyth, Dan, et al.
Pubblicazione: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
di: Dong, Lukuang, et al.
Pubblicazione: (2026)
di: Dong, Lukuang, et al.
Pubblicazione: (2026)
Predicting speech intelligibility in older adults for speech enhancement using the Gammachirp Envelope Similarity Index, GESI
di: Yamamoto, Ayako, et al.
Pubblicazione: (2025)
di: Yamamoto, Ayako, et al.
Pubblicazione: (2025)
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
di: Mack, Wolfgang, et al.
Pubblicazione: (2025)
di: Mack, Wolfgang, et al.
Pubblicazione: (2025)
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
di: Terashima, Ryo, et al.
Pubblicazione: (2025)
di: Terashima, Ryo, et al.
Pubblicazione: (2025)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
di: Igarashi, Takuto, et al.
Pubblicazione: (2024)
di: Igarashi, Takuto, et al.
Pubblicazione: (2024)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
di: Saito, Yuki, et al.
Pubblicazione: (2024)
di: Saito, Yuki, et al.
Pubblicazione: (2024)
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
di: Malisz, Zofia, et al.
Pubblicazione: (2024)
di: Malisz, Zofia, et al.
Pubblicazione: (2024)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
di: Ohnaka, Hien, et al.
Pubblicazione: (2026)
di: Ohnaka, Hien, et al.
Pubblicazione: (2026)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
di: Li, Xuyuan, et al.
Pubblicazione: (2023)
di: Li, Xuyuan, et al.
Pubblicazione: (2023)
Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech
di: Jiang, Pan-Pan, et al.
Pubblicazione: (2024)
di: Jiang, Pan-Pan, et al.
Pubblicazione: (2024)
Transferable speech-to-text large language model alignment module
di: Wu, Boyong, et al.
Pubblicazione: (2024)
di: Wu, Boyong, et al.
Pubblicazione: (2024)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
di: Huang, Ziling, et al.
Pubblicazione: (2025)
di: Huang, Ziling, et al.
Pubblicazione: (2025)
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
di: Saon, George, et al.
Pubblicazione: (2025)
di: Saon, George, et al.
Pubblicazione: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
Disentangling peripheral hearing loss from central and cognitive effects on speech intelligibility in older adults
di: Irino, Toshio, et al.
Pubblicazione: (2025)
di: Irino, Toshio, et al.
Pubblicazione: (2025)
Universal Score-based Speech Enhancement with High Content Preservation
di: Scheibler, Robin, et al.
Pubblicazione: (2024)
di: Scheibler, Robin, et al.
Pubblicazione: (2024)
Good practices for evaluation of synthesized speech
di: Cooper, Erica, et al.
Pubblicazione: (2025)
di: Cooper, Erica, et al.
Pubblicazione: (2025)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
di: Jeong, Jaeseok, et al.
Pubblicazione: (2025)
di: Jeong, Jaeseok, et al.
Pubblicazione: (2025)
Sound event detection with audio-text models and heterogeneous temporal annotations
di: Harju, Manu, et al.
Pubblicazione: (2025)
di: Harju, Manu, et al.
Pubblicazione: (2025)
Prominence-aware automatic speech recognition for conversational speech
di: Linke, Julian, et al.
Pubblicazione: (2025)
di: Linke, Julian, et al.
Pubblicazione: (2025)
Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation
di: Ma, Jianbo, et al.
Pubblicazione: (2026)
di: Ma, Jianbo, et al.
Pubblicazione: (2026)
An efficient text augmentation approach for contextualized Mandarin speech recognition
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
Probing mental health information in speech foundation models
di: de Gennes, Marc, et al.
Pubblicazione: (2024)
di: de Gennes, Marc, et al.
Pubblicazione: (2024)
WhisperFlow: speech foundation models in real time
di: Wang, Rongxiang, et al.
Pubblicazione: (2024)
di: Wang, Rongxiang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
di: Ohnaka, Hien, et al.
Pubblicazione: (2025) -
CC-G2PnP: Streaming Grapheme-to-Phoneme and prosody with Conformer-CTC for unsegmented languages
di: Shirahata, Yuma, et al.
Pubblicazione: (2026) -
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
di: Yamamoto, Ryuichi, et al.
Pubblicazione: (2024) -
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
di: Kawamura, Masaya, et al.
Pubblicazione: (2024) -
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
di: Ma, Te, et al.
Pubblicazione: (2025)