A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Onda, Kentaro, Park, Joonyong, Minematsu, Nobuaki, Saito, Daisuke |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
par: Onda, Kentaro, et autres
Publié: (2025)
par: Onda, Kentaro, et autres
Publié: (2025)
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
par: Park, Joonyong, et autres
Publié: (2024)
par: Park, Joonyong, et autres
Publié: (2024)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
par: Onda, Kentaro, et autres
Publié: (2025)
par: Onda, Kentaro, et autres
Publié: (2025)
Benchmarking Prosody Encoding in Discrete Speech Tokens
par: Onda, Kentaro, et autres
Publié: (2025)
par: Onda, Kentaro, et autres
Publié: (2025)
A Pilot Study of Applying Sequence-to-Sequence Voice Conversion to Evaluate the Intelligibility of L2 Speech Using a Native Speaker's Shadowings
par: Geng, Haopeng, et autres
Publié: (2024)
par: Geng, Haopeng, et autres
Publié: (2024)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
par: Geng, Haopeng, et autres
Publié: (2024)
par: Geng, Haopeng, et autres
Publié: (2024)
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
par: Park, Joonyong, et autres
Publié: (2025)
par: Park, Joonyong, et autres
Publié: (2025)
A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
par: Geng, Haopeng, et autres
Publié: (2025)
par: Geng, Haopeng, et autres
Publié: (2025)
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
par: Huang, Zhijie, et autres
Publié: (2026)
par: Huang, Zhijie, et autres
Publié: (2026)
AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style
par: Park, Joonyong, et autres
Publié: (2026)
par: Park, Joonyong, et autres
Publié: (2026)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
par: Sun, Haitong, et autres
Publié: (2026)
par: Sun, Haitong, et autres
Publié: (2026)
Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis
par: Geng, Haopeng, et autres
Publié: (2026)
par: Geng, Haopeng, et autres
Publié: (2026)
CAMEO: Collection of Multilingual Emotional Speech Corpora
par: Christop, Iwona, et autres
Publié: (2025)
par: Christop, Iwona, et autres
Publié: (2025)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
par: Chen, Peikun, et autres
Publié: (2024)
par: Chen, Peikun, et autres
Publié: (2024)
Differentiable K-means for Fully-optimized Discrete Token-based ASR
par: Onda, Kentaro, et autres
Publié: (2025)
par: Onda, Kentaro, et autres
Publié: (2025)
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
par: Xu, Jing, et autres
Publié: (2024)
par: Xu, Jing, et autres
Publié: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
par: Peng, Yifan, et autres
Publié: (2024)
par: Peng, Yifan, et autres
Publié: (2024)
Automatic Speech Recognition for Non-Native English: Accuracy and Disfluency Handling
par: McGuire, Michael
Publié: (2025)
par: McGuire, Michael
Publié: (2025)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
par: Masuyama, Yoshiki, et autres
Publié: (2024)
par: Masuyama, Yoshiki, et autres
Publié: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
par: Nespoli, Francesco, et autres
Publié: (2024)
par: Nespoli, Francesco, et autres
Publié: (2024)
LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech
par: Kim, Haechan, et autres
Publié: (2024)
par: Kim, Haechan, et autres
Publié: (2024)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
par: Deng, Keqi, et autres
Publié: (2025)
par: Deng, Keqi, et autres
Publié: (2025)
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
par: Wills, Simone, et autres
Publié: (2023)
par: Wills, Simone, et autres
Publié: (2023)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
par: Park, Chanho, et autres
Publié: (2023)
par: Park, Chanho, et autres
Publié: (2023)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
par: Suda, Hitoshi, et autres
Publié: (2025)
par: Suda, Hitoshi, et autres
Publié: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
par: Yamauchi, Kazuki, et autres
Publié: (2024)
par: Yamauchi, Kazuki, et autres
Publié: (2024)
Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach
par: Abdullah, Abdulhady Abas, et autres
Publié: (2024)
par: Abdullah, Abdulhady Abas, et autres
Publié: (2024)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
par: Nakata, Wataru, et autres
Publié: (2024)
par: Nakata, Wataru, et autres
Publié: (2024)
Speaker-Aware Simulation Improves Conversational Speech Recognition
par: Gedeon, Máté, et autres
Publié: (2026)
par: Gedeon, Máté, et autres
Publié: (2026)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
par: Igarashi, Takuto, et autres
Publié: (2024)
par: Igarashi, Takuto, et autres
Publié: (2024)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
par: Lee, Jeehyun, et autres
Publié: (2024)
par: Lee, Jeehyun, et autres
Publié: (2024)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
par: Wang, Yuejiao, et autres
Publié: (2024)
par: Wang, Yuejiao, et autres
Publié: (2024)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
par: Peng, Yifan, et autres
Publié: (2024)
par: Peng, Yifan, et autres
Publié: (2024)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
par: Kong, Jungil, et autres
Publié: (2023)
par: Kong, Jungil, et autres
Publié: (2023)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
par: Liu, Henglyu, et autres
Publié: (2025)
par: Liu, Henglyu, et autres
Publié: (2025)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
par: Yamamoto, Ryuichi, et autres
Publié: (2024)
par: Yamamoto, Ryuichi, et autres
Publié: (2024)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
par: Wang, Hankun, et autres
Publié: (2024)
par: Wang, Hankun, et autres
Publié: (2024)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
par: Park, Chanho, et autres
Publié: (2024)
par: Park, Chanho, et autres
Publié: (2024)
Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
par: Park, ChaeHun, et autres
Publié: (2024)
par: Park, ChaeHun, et autres
Publié: (2024)
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
par: Choi, Anna Seo Gyeong, et autres
Publié: (2025)
par: Choi, Anna Seo Gyeong, et autres
Publié: (2025)
Documents similaires
-
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
par: Onda, Kentaro, et autres
Publié: (2025) -
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
par: Park, Joonyong, et autres
Publié: (2024) -
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
par: Onda, Kentaro, et autres
Publié: (2025) -
Benchmarking Prosody Encoding in Discrete Speech Tokens
par: Onda, Kentaro, et autres
Publié: (2025) -
A Pilot Study of Applying Sequence-to-Sequence Voice Conversion to Evaluate the Intelligibility of L2 Speech Using a Native Speaker's Shadowings
par: Geng, Haopeng, et autres
Publié: (2024)