Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chang, Kalvin, Shao, Yiwen, Li, Jiahong, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Phonotactic Complexity across Dialects
von: Shim, Ryan Soh-Eun, et al.
Veröffentlicht: (2024)
von: Shim, Ryan Soh-Eun, et al.
Veröffentlicht: (2024)
AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning
von: Shao, Yiwen, et al.
Veröffentlicht: (2025)
von: Shao, Yiwen, et al.
Veröffentlicht: (2025)
Dolphin-CN-Dialect: Where Chinese Dialects Matter
von: Meng, Yangyang, et al.
Veröffentlicht: (2026)
von: Meng, Yangyang, et al.
Veröffentlicht: (2026)
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
Efficient Multilingual ASR Finetuning via LoRA Language Experts
von: Li, Jiahong, et al.
Veröffentlicht: (2025)
von: Li, Jiahong, et al.
Veröffentlicht: (2025)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
von: Blaschke, Verena, et al.
Veröffentlicht: (2025)
von: Blaschke, Verena, et al.
Veröffentlicht: (2025)
Phonetic Modeling of Dialectal Variation in Vietnamese Speech
von: Hoang, Quan Ngoc, et al.
Veröffentlicht: (2026)
von: Hoang, Quan Ngoc, et al.
Veröffentlicht: (2026)
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models
von: Khamis, Ahmed Khaled, et al.
Veröffentlicht: (2026)
von: Khamis, Ahmed Khaled, et al.
Veröffentlicht: (2026)
RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects
von: Hassan, Md. Rezuwan, et al.
Veröffentlicht: (2025)
von: Hassan, Md. Rezuwan, et al.
Veröffentlicht: (2025)
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology
von: Sullivan, Peter, et al.
Veröffentlicht: (2026)
von: Sullivan, Peter, et al.
Veröffentlicht: (2026)
Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translation
von: Park, Keunhyeung, et al.
Veröffentlicht: (2025)
von: Park, Keunhyeung, et al.
Veröffentlicht: (2025)
BanglaTalk: Towards Real-Time Speech Assistance for Bengali Regional Dialects
von: Hasan, Jakir, et al.
Veröffentlicht: (2025)
von: Hasan, Jakir, et al.
Veröffentlicht: (2025)
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
von: Blaschke, Verena, et al.
Veröffentlicht: (2025)
von: Blaschke, Verena, et al.
Veröffentlicht: (2025)
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
von: Huo, Mingyue, et al.
Veröffentlicht: (2026)
von: Huo, Mingyue, et al.
Veröffentlicht: (2026)
WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
von: Oberkircher, Lena S., et al.
Veröffentlicht: (2026)
von: Oberkircher, Lena S., et al.
Veröffentlicht: (2026)
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models
von: Bafna, Niyati, et al.
Veröffentlicht: (2025)
von: Bafna, Niyati, et al.
Veröffentlicht: (2025)
What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects
von: Blaschke, Verena, et al.
Veröffentlicht: (2024)
von: Blaschke, Verena, et al.
Veröffentlicht: (2024)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
Towards Comprehensive Detection of Chinese Harmful Memes
von: Lu, Junyu, et al.
Veröffentlicht: (2024)
von: Lu, Junyu, et al.
Veröffentlicht: (2024)
Dialectal Coverage And Generalization in Arabic Speech Recognition
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2024)
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2024)
Linear Semantic Segmentation for Low-Resource Spoken Dialects
von: Chirkunov, Kirill, et al.
Veröffentlicht: (2026)
von: Chirkunov, Kirill, et al.
Veröffentlicht: (2026)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
von: Dorn, Rebecca, et al.
Veröffentlicht: (2024)
von: Dorn, Rebecca, et al.
Veröffentlicht: (2024)
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2021)
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2021)
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Picturized and Recited with Dialects: A Multimodal Chinese Representation Framework for Sentiment Analysis of Classical Chinese Poetry
von: Du, Xiaocong, et al.
Veröffentlicht: (2025)
von: Du, Xiaocong, et al.
Veröffentlicht: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
ArFake: A Multi-Dialect Benchmark and Baselines for Arabic Spoof-Speech Detection
von: Maged, Mohamed, et al.
Veröffentlicht: (2025)
von: Maged, Mohamed, et al.
Veröffentlicht: (2025)
TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Exploiting Dialect Identification in Automatic Dialectal Text Normalization
von: Alhafni, Bashar, et al.
Veröffentlicht: (2024)
von: Alhafni, Bashar, et al.
Veröffentlicht: (2024)
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
von: Shim, Ryan Soh-Eun, et al.
Veröffentlicht: (2026)
von: Shim, Ryan Soh-Eun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Phonotactic Complexity across Dialects
von: Shim, Ryan Soh-Eun, et al.
Veröffentlicht: (2024) -
AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning
von: Shao, Yiwen, et al.
Veröffentlicht: (2025) -
Dolphin-CN-Dialect: Where Chinese Dialects Matter
von: Meng, Yangyang, et al.
Veröffentlicht: (2026) -
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
von: Di, Xinhan, et al.
Veröffentlicht: (2024) -
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)