Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Joonyong, Saito, Daisuke, Minematsu, Nobuaki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
von: Onda, Kentaro, et al.
Veröffentlicht: (2024)
von: Onda, Kentaro, et al.
Veröffentlicht: (2024)
A Pilot Study of Applying Sequence-to-Sequence Voice Conversion to Evaluate the Intelligibility of L2 Speech Using a Native Speaker's Shadowings
von: Geng, Haopeng, et al.
Veröffentlicht: (2024)
von: Geng, Haopeng, et al.
Veröffentlicht: (2024)
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
von: Geng, Haopeng, et al.
Veröffentlicht: (2024)
von: Geng, Haopeng, et al.
Veröffentlicht: (2024)
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
von: Huang, Zhijie, et al.
Veröffentlicht: (2026)
von: Huang, Zhijie, et al.
Veröffentlicht: (2026)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
von: Geng, Haopeng, et al.
Veröffentlicht: (2025)
von: Geng, Haopeng, et al.
Veröffentlicht: (2025)
AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style
von: Park, Joonyong, et al.
Veröffentlicht: (2026)
von: Park, Joonyong, et al.
Veröffentlicht: (2026)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis
von: Geng, Haopeng, et al.
Veröffentlicht: (2026)
von: Geng, Haopeng, et al.
Veröffentlicht: (2026)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
von: Park, Chanho, et al.
Veröffentlicht: (2023)
von: Park, Chanho, et al.
Veröffentlicht: (2023)
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
Is Self-Supervised Learning Enough to Fill in the Gap? A Study on Speech Inpainting
von: Asaad, Ihab, et al.
Veröffentlicht: (2024)
von: Asaad, Ihab, et al.
Veröffentlicht: (2024)
Interface Design for Self-Supervised Speech Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
von: Jang, Kangwook, et al.
Veröffentlicht: (2023)
von: Jang, Kangwook, et al.
Veröffentlicht: (2023)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
von: Phuong, Tuan Dat, et al.
Veröffentlicht: (2025)
von: Phuong, Tuan Dat, et al.
Veröffentlicht: (2025)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2025)
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2025)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Self-Powered LLM Modality Expansion for Large Speech-Text Models
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition
von: Saliba, Alexandra, et al.
Veröffentlicht: (2024)
von: Saliba, Alexandra, et al.
Veröffentlicht: (2024)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
von: Manakul, Potsawee, et al.
Veröffentlicht: (2025)
von: Manakul, Potsawee, et al.
Veröffentlicht: (2025)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Boosting Large Language Model for Speech Synthesis: An Empirical Study
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
von: Park, Joonyong, et al.
Veröffentlicht: (2025) -
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
von: Onda, Kentaro, et al.
Veröffentlicht: (2024) -
A Pilot Study of Applying Sequence-to-Sequence Voice Conversion to Evaluate the Intelligibility of L2 Speech Using a Native Speaker's Shadowings
von: Geng, Haopeng, et al.
Veröffentlicht: (2024) -
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025) -
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
von: Geng, Haopeng, et al.
Veröffentlicht: (2024)