Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yushen, Liu, Junzhe, Tu, Yujie, Niu, Zhikang, Liang, Yuzhe, Qiang, Chunyu, Zhang, Chen, Yu, Kai, Chen, Xie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
von: Zheng, Qixi, et al.
Veröffentlicht: (2025)
von: Zheng, Qixi, et al.
Veröffentlicht: (2025)
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
von: Chen, Yushen, et al.
Veröffentlicht: (2024)
von: Chen, Yushen, et al.
Veröffentlicht: (2024)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
Dialectal Coverage And Generalization in Arabic Speech Recognition
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2024)
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2024)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
von: Niu, Zhikang, et al.
Veröffentlicht: (2024)
von: Niu, Zhikang, et al.
Veröffentlicht: (2024)
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
von: Chen, Yushen, et al.
Veröffentlicht: (2025)
von: Chen, Yushen, et al.
Veröffentlicht: (2025)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
von: Niu, Rui, et al.
Veröffentlicht: (2025)
von: Niu, Rui, et al.
Veröffentlicht: (2025)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
von: Xu, Kai-Tuo, et al.
Veröffentlicht: (2025)
von: Xu, Kai-Tuo, et al.
Veröffentlicht: (2025)
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
von: Naouara, Hedi, et al.
Veröffentlicht: (2025)
von: Naouara, Hedi, et al.
Veröffentlicht: (2025)
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
Towards Flow-Matching-based TTS without Classifier-Free Guidance
von: Liang, Yuzhe, et al.
Veröffentlicht: (2025)
von: Liang, Yuzhe, et al.
Veröffentlicht: (2025)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Frequency-mix Knowledge Distillation for Fake Speech Detection
von: Fan, Cunhang, et al.
Veröffentlicht: (2024)
von: Fan, Cunhang, et al.
Veröffentlicht: (2024)
A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR
von: Biswas, Swadhin, et al.
Veröffentlicht: (2025)
von: Biswas, Swadhin, et al.
Veröffentlicht: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
von: Zheng, Qixi, et al.
Veröffentlicht: (2025) -
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
von: Chen, Yushen, et al.
Veröffentlicht: (2024) -
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025) -
Dialectal Coverage And Generalization in Arabic Speech Recognition
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2024) -
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)