Rethinking Discrete Speech Representation Tokens for Accent Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhong, Jinzuomu, Wang, Yi, Richmond, Korin, Bell, Peter |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pairwise Evaluation of Accent Similarity in Speech Synthesis
por: Zhong, Jinzuomu, et al.
Publicado: (2025)
por: Zhong, Jinzuomu, et al.
Publicado: (2025)
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
por: Zhong, Jinzuomu, et al.
Publicado: (2024)
por: Zhong, Jinzuomu, et al.
Publicado: (2024)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
por: Sun, Siqi, et al.
Publicado: (2024)
por: Sun, Siqi, et al.
Publicado: (2024)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
por: Sun, Yujia, et al.
Publicado: (2024)
por: Sun, Yujia, et al.
Publicado: (2024)
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
por: Zhong, Jinzuomu, et al.
Publicado: (2023)
por: Zhong, Jinzuomu, et al.
Publicado: (2023)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
por: Sanders, Nicholas, et al.
Publicado: (2025)
por: Sanders, Nicholas, et al.
Publicado: (2025)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
por: Gong, Cheng, et al.
Publicado: (2023)
por: Gong, Cheng, et al.
Publicado: (2023)
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
por: Bafna, Niyati, et al.
Publicado: (2025)
por: Bafna, Niyati, et al.
Publicado: (2025)
Crossmodal ASR Error Correction with Discrete Speech Units
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
por: Venkateswaran, Nitin, et al.
Publicado: (2025)
por: Venkateswaran, Nitin, et al.
Publicado: (2025)
Benchmarking Prosody Encoding in Discrete Speech Tokens
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
Performant ASR Models for Medical Entities in Accented Speech
por: Afonja, Tejumade, et al.
Publicado: (2024)
por: Afonja, Tejumade, et al.
Publicado: (2024)
Children's Speech Recognition through Discrete Token Enhancement
por: Sukhadia, Vrunda N., et al.
Publicado: (2024)
por: Sukhadia, Vrunda N., et al.
Publicado: (2024)
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
por: Sasu, David, et al.
Publicado: (2025)
por: Sasu, David, et al.
Publicado: (2025)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
por: Do, Cong-Thanh, et al.
Publicado: (2024)
por: Do, Cong-Thanh, et al.
Publicado: (2024)
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents
por: Owodunni, Abraham Toluwase, et al.
Publicado: (2024)
por: Owodunni, Abraham Toluwase, et al.
Publicado: (2024)
Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
por: Sameti, Mohammad Hossein, et al.
Publicado: (2025)
por: Sameti, Mohammad Hossein, et al.
Publicado: (2025)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
por: Wang, Dingdong, et al.
Publicado: (2024)
por: Wang, Dingdong, et al.
Publicado: (2024)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
por: Zhang, Xin, et al.
Publicado: (2023)
por: Zhang, Xin, et al.
Publicado: (2023)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
por: Osakuade, Opeyemi, et al.
Publicado: (2024)
por: Osakuade, Opeyemi, et al.
Publicado: (2024)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
Advancing African-Accented Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models
por: Dossou, Bonaventure F. P.
Publicado: (2023)
por: Dossou, Bonaventure F. P.
Publicado: (2023)
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
por: Minixhofer, Christoph, et al.
Publicado: (2025)
por: Minixhofer, Christoph, et al.
Publicado: (2025)
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
Continuous Speech Tokenizer in Text To Speech
por: Li, Yixing, et al.
Publicado: (2024)
por: Li, Yixing, et al.
Publicado: (2024)
Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
por: Bai, Qibing, et al.
Publicado: (2025)
por: Bai, Qibing, et al.
Publicado: (2025)
Exploring SSL Discrete Tokens for Multilingual ASR
por: Cui, Mingyu, et al.
Publicado: (2024)
por: Cui, Mingyu, et al.
Publicado: (2024)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
por: Tseng, Liang-Hsuan, et al.
Publicado: (2025)
por: Tseng, Liang-Hsuan, et al.
Publicado: (2025)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
Frontend Token Enhancement for Token-Based Speech Recognition
por: Ashihara, Takanori, et al.
Publicado: (2026)
por: Ashihara, Takanori, et al.
Publicado: (2026)
Next Tokens Denoising for Speech Synthesis
por: Liu, Yanqing, et al.
Publicado: (2025)
por: Liu, Yanqing, et al.
Publicado: (2025)
Factorized RVQ-GAN For Disentangled Speech Tokenization
por: Khurana, Sameer, et al.
Publicado: (2025)
por: Khurana, Sameer, et al.
Publicado: (2025)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
por: Shon, Suwon, et al.
Publicado: (2024)
por: Shon, Suwon, et al.
Publicado: (2024)
Exploring the Benefits of Tokenization of Discrete Acoustic Units
por: Dekel, Avihu, et al.
Publicado: (2024)
por: Dekel, Avihu, et al.
Publicado: (2024)
STAB: Speech Tokenizer Assessment Benchmark
por: Vashishth, Shikhar, et al.
Publicado: (2024)
por: Vashishth, Shikhar, et al.
Publicado: (2024)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
por: Han, Zhichen, et al.
Publicado: (2024)
por: Han, Zhichen, et al.
Publicado: (2024)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
por: Nguyen, Tuan-Nam, et al.
Publicado: (2025)
por: Nguyen, Tuan-Nam, et al.
Publicado: (2025)
Compact Speech Translation Models via Discrete Speech Units Pretraining
por: Lam, Tsz Kin, et al.
Publicado: (2024)
por: Lam, Tsz Kin, et al.
Publicado: (2024)
LAST: Language Model Aware Speech Tokenization
por: Turetzky, Arnon, et al.
Publicado: (2024)
por: Turetzky, Arnon, et al.
Publicado: (2024)
Ejemplares similares
-
Pairwise Evaluation of Accent Similarity in Speech Synthesis
por: Zhong, Jinzuomu, et al.
Publicado: (2025) -
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
por: Zhong, Jinzuomu, et al.
Publicado: (2024) -
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
por: Sun, Siqi, et al.
Publicado: (2024) -
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
por: Sun, Yujia, et al.
Publicado: (2024) -
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
por: Zhong, Jinzuomu, et al.
Publicado: (2023)