MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Joonyong, Saito, Daisuke, Minematsu, Nobuaki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
by: Park, Joonyong, et al.
Published: (2024)
by: Park, Joonyong, et al.
Published: (2024)
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
by: Onda, Kentaro, et al.
Published: (2024)
by: Onda, Kentaro, et al.
Published: (2024)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
by: Geng, Haopeng, et al.
Published: (2024)
by: Geng, Haopeng, et al.
Published: (2024)
A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
by: Geng, Haopeng, et al.
Published: (2025)
by: Geng, Haopeng, et al.
Published: (2025)
EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
by: Park, Joonyong, et al.
Published: (2025)
by: Park, Joonyong, et al.
Published: (2025)
A Pilot Study of Applying Sequence-to-Sequence Voice Conversion to Evaluate the Intelligibility of L2 Speech Using a Native Speaker's Shadowings
by: Geng, Haopeng, et al.
Published: (2024)
by: Geng, Haopeng, et al.
Published: (2024)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
by: Onda, Kentaro, et al.
Published: (2025)
by: Onda, Kentaro, et al.
Published: (2025)
Benchmarking Prosody Encoding in Discrete Speech Tokens
by: Onda, Kentaro, et al.
Published: (2025)
by: Onda, Kentaro, et al.
Published: (2025)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
by: Onda, Kentaro, et al.
Published: (2025)
by: Onda, Kentaro, et al.
Published: (2025)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
by: Sun, Haitong, et al.
Published: (2026)
by: Sun, Haitong, et al.
Published: (2026)
AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style
by: Park, Joonyong, et al.
Published: (2026)
by: Park, Joonyong, et al.
Published: (2026)
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
by: Huang, Zhijie, et al.
Published: (2026)
by: Huang, Zhijie, et al.
Published: (2026)
Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis
by: Geng, Haopeng, et al.
Published: (2026)
by: Geng, Haopeng, et al.
Published: (2026)
AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup
by: Carvalho, Carlos, et al.
Published: (2024)
by: Carvalho, Carlos, et al.
Published: (2024)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
by: Cohen, Eyal, et al.
Published: (2025)
by: Cohen, Eyal, et al.
Published: (2025)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
by: Shah, Neil, et al.
Published: (2024)
by: Shah, Neil, et al.
Published: (2024)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
by: Tseng, Wei-Cheng, et al.
Published: (2025)
by: Tseng, Wei-Cheng, et al.
Published: (2025)
Rethinking Mamba in Speech Processing by Self-Supervised Models
by: Zhang, Xiangyu, et al.
Published: (2024)
by: Zhang, Xiangyu, et al.
Published: (2024)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
by: Menon, Aditya Srinivas, et al.
Published: (2026)
by: Menon, Aditya Srinivas, et al.
Published: (2026)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
by: Chung, Raymond
Published: (2026)
by: Chung, Raymond
Published: (2026)
AS-Speech: Adaptive Style For Speech Synthesis
by: Li, Zhipeng, et al.
Published: (2024)
by: Li, Zhipeng, et al.
Published: (2024)
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
by: Guo, Xin, et al.
Published: (2026)
by: Guo, Xin, et al.
Published: (2026)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
by: Li, Hengyu, et al.
Published: (2024)
by: Li, Hengyu, et al.
Published: (2024)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
by: Aldeneh, Zakaria, et al.
Published: (2024)
by: Aldeneh, Zakaria, et al.
Published: (2024)
Analysing the Language of Neural Audio Codecs
by: Park, Joonyong, et al.
Published: (2025)
by: Park, Joonyong, et al.
Published: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
by: Du, Chenpeng, et al.
Published: (2022)
by: Du, Chenpeng, et al.
Published: (2022)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
by: Li, Yinghao Aaron, et al.
Published: (2025)
by: Li, Yinghao Aaron, et al.
Published: (2025)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
by: Kuhlmann, Michael, et al.
Published: (2026)
by: Kuhlmann, Michael, et al.
Published: (2026)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
by: Ogg, Mattson, et al.
Published: (2025)
by: Ogg, Mattson, et al.
Published: (2025)
Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection
by: Hanilçi, Cemal, et al.
Published: (2026)
by: Hanilçi, Cemal, et al.
Published: (2026)
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
by: C, Shiva Kumar, et al.
Published: (2025)
by: C, Shiva Kumar, et al.
Published: (2025)
Ring Mixing with Auxiliary Signal-to-Consistency-Error Ratio Loss for Unsupervised Denoising in Speech Separation
by: Maciejewski, Matthew, et al.
Published: (2026)
by: Maciejewski, Matthew, et al.
Published: (2026)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
by: Tsunoo, Emiru, et al.
Published: (2024)
by: Tsunoo, Emiru, et al.
Published: (2024)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
by: Zhao, Jinzheng, et al.
Published: (2024)
by: Zhao, Jinzheng, et al.
Published: (2024)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
by: Farhadipour, Aref, et al.
Published: (2024)
by: Farhadipour, Aref, et al.
Published: (2024)
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
by: Kutsakov, Aleksandr, et al.
Published: (2025)
by: Kutsakov, Aleksandr, et al.
Published: (2025)
Similar Items
-
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
by: Park, Joonyong, et al.
Published: (2024) -
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
by: Onda, Kentaro, et al.
Published: (2024) -
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
by: Geng, Haopeng, et al.
Published: (2024) -
A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
by: Geng, Haopeng, et al.
Published: (2025) -
EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
by: Park, Joonyong, et al.
Published: (2025)