Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ohnaka, Hien, Shirahata, Yuma, Park, Byeongseon, Yamamoto, Ryuichi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CC-G2PnP: Streaming Grapheme-to-Phoneme and prosody with Conformer-CTC for unsegmented languages
von: Shirahata, Yuma, et al.
Veröffentlicht: (2026)
von: Shirahata, Yuma, et al.
Veröffentlicht: (2026)
Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech
von: Kolani, Yakov, et al.
Veröffentlicht: (2025)
von: Kolani, Yakov, et al.
Veröffentlicht: (2025)
Multitask Learning for Grapheme-to-Phoneme Conversion of Anglicisms in German Speech Recognition
von: Pritzen, Julia, et al.
Veröffentlicht: (2021)
von: Pritzen, Julia, et al.
Veröffentlicht: (2021)
Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
von: Ohnaka, Hien, et al.
Veröffentlicht: (2026)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2026)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
GraphemeAug: A Systematic Approach to Synthesized Hard Negative Keyword Spotting Examples
von: Zhang, Harry, et al.
Veröffentlicht: (2025)
von: Zhang, Harry, et al.
Veröffentlicht: (2025)
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
von: Kawamura, Masaya, et al.
Veröffentlicht: (2024)
von: Kawamura, Masaya, et al.
Veröffentlicht: (2024)
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
von: Shirahata, Yuma, et al.
Veröffentlicht: (2024)
von: Shirahata, Yuma, et al.
Veröffentlicht: (2024)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
von: Kawamura, Masaya, et al.
Veröffentlicht: (2025)
PSST! Prosodic Speech Segmentation with Transformers
von: Roll, Nathan, et al.
Veröffentlicht: (2023)
von: Roll, Nathan, et al.
Veröffentlicht: (2023)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
Unsupervised speech enhancement with spectral kurtosis and double deep priors
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2024)
Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning
von: Woszczyk, Dominika, et al.
Veröffentlicht: (2025)
von: Woszczyk, Dominika, et al.
Veröffentlicht: (2025)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
Universal Score-based Speech Enhancement with High Content Preservation
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
Which Prosodic Features Matter Most for Pragmatics?
von: Ward, Nigel G., et al.
Veröffentlicht: (2024)
von: Ward, Nigel G., et al.
Veröffentlicht: (2024)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
von: Gállego, Gerard I., et al.
Veröffentlicht: (2025)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2025)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
von: Dong, Lukuan, et al.
Veröffentlicht: (2024)
von: Dong, Lukuan, et al.
Veröffentlicht: (2024)
A Functional Trade-off between Prosodic and Semantic Cues in Conveying Sarcasm
von: Li, Zhu, et al.
Veröffentlicht: (2024)
von: Li, Zhu, et al.
Veröffentlicht: (2024)
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
von: Azad, Asif, et al.
Veröffentlicht: (2026)
von: Azad, Asif, et al.
Veröffentlicht: (2026)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration
von: Karita, Shigeki, et al.
Veröffentlicht: (2025)
von: Karita, Shigeki, et al.
Veröffentlicht: (2025)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
von: Poli, Maxime, et al.
Veröffentlicht: (2024)
von: Poli, Maxime, et al.
Veröffentlicht: (2024)
Cross-Utterance Conditioned VAE for Speech Generation
von: Li, Yang, et al.
Veröffentlicht: (2023)
von: Li, Yang, et al.
Veröffentlicht: (2023)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
von: Liu, Sen, et al.
Veröffentlicht: (2024)
von: Liu, Sen, et al.
Veröffentlicht: (2024)
Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
von: Fetrat, Mahta, et al.
Veröffentlicht: (2025)
von: Fetrat, Mahta, et al.
Veröffentlicht: (2025)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
von: Chary, Podakanti Satyajith
Veröffentlicht: (2024)
von: Chary, Podakanti Satyajith
Veröffentlicht: (2024)
Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CC-G2PnP: Streaming Grapheme-to-Phoneme and prosody with Conformer-CTC for unsegmented languages
von: Shirahata, Yuma, et al.
Veröffentlicht: (2026) -
Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech
von: Kolani, Yakov, et al.
Veröffentlicht: (2025) -
Multitask Learning for Grapheme-to-Phoneme Conversion of Anglicisms in German Speech Recognition
von: Pritzen, Julia, et al.
Veröffentlicht: (2021) -
Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
von: Ohnaka, Hien, et al.
Veröffentlicht: (2026) -
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)