Guardado en:
| Autores principales: | Khamis, Ahmed Khaled, Ali, Hesham |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.15675 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GATech at AbjadGenEval Shared Task: Multilingual Embeddings for Arabic Machine-Generated Text Classification
por: Khamis, Ahmed Khaled
Publicado: (2026)
por: Khamis, Ahmed Khaled
Publicado: (2026)
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
por: Dua, Karan, et al.
Publicado: (2025)
por: Dua, Karan, et al.
Publicado: (2025)
GATech at AbjadMed: Bidirectional Encoders vs. Causal Decoders: Insights from 82-Class Arabic Medical Classification
por: Khamis, Ahmed Khaled
Publicado: (2026)
por: Khamis, Ahmed Khaled
Publicado: (2026)
Phonetic Modeling of Dialectal Variation in Vietnamese Speech
por: Hoang, Quan Ngoc, et al.
Publicado: (2026)
por: Hoang, Quan Ngoc, et al.
Publicado: (2026)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
por: Doan, Khai Duy, et al.
Publicado: (2024)
por: Doan, Khai Duy, et al.
Publicado: (2024)
RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects
por: Hassan, Md. Rezuwan, et al.
Publicado: (2025)
por: Hassan, Md. Rezuwan, et al.
Publicado: (2025)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
por: Xu, Tianyi, et al.
Publicado: (2025)
por: Xu, Tianyi, et al.
Publicado: (2025)
Computational Linguistics Meets Libyan Dialect: A Study on Dialect Identification
por: Essgaer, Mansour, et al.
Publicado: (2025)
por: Essgaer, Mansour, et al.
Publicado: (2025)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
por: Rossenbach, Nick, et al.
Publicado: (2024)
por: Rossenbach, Nick, et al.
Publicado: (2024)
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
por: Ahmed, Ahmed Haj, et al.
Publicado: (2024)
por: Ahmed, Ahmed Haj, et al.
Publicado: (2024)
Scaling Arabic Medical Chatbots Using Synthetic Data: Enhancing Generative AI with Synthetic Patient Records
por: Allam, Abdulrahman, et al.
Publicado: (2025)
por: Allam, Abdulrahman, et al.
Publicado: (2025)
A Dialectic Pipeline for Improving LLM Robustness
por: Candussio, Sara
Publicado: (2026)
por: Candussio, Sara
Publicado: (2026)
A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
por: Blaschke, Verena, et al.
Publicado: (2025)
por: Blaschke, Verena, et al.
Publicado: (2025)
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
por: Blaschke, Verena, et al.
Publicado: (2025)
por: Blaschke, Verena, et al.
Publicado: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
por: Allbert, Rumi, et al.
Publicado: (2025)
por: Allbert, Rumi, et al.
Publicado: (2025)
ArFake: A Multi-Dialect Benchmark and Baselines for Arabic Spoof-Speech Detection
por: Maged, Mohamed, et al.
Publicado: (2025)
por: Maged, Mohamed, et al.
Publicado: (2025)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
por: Zeng, Aohan, et al.
Publicado: (2024)
por: Zeng, Aohan, et al.
Publicado: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
por: Futami, Hayato, et al.
Publicado: (2025)
por: Futami, Hayato, et al.
Publicado: (2025)
Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects
por: Chang, Kalvin, et al.
Publicado: (2026)
por: Chang, Kalvin, et al.
Publicado: (2026)
TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation
por: Liu, Yutong, et al.
Publicado: (2025)
por: Liu, Yutong, et al.
Publicado: (2025)
PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-Speech
por: Wong, Michel, et al.
Publicado: (2025)
por: Wong, Michel, et al.
Publicado: (2025)
Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages
por: Popescu-Belis, Andrei, et al.
Publicado: (2025)
por: Popescu-Belis, Andrei, et al.
Publicado: (2025)
WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
por: Dai, Yuhang, et al.
Publicado: (2025)
por: Dai, Yuhang, et al.
Publicado: (2025)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
por: Dorn, Rebecca, et al.
Publicado: (2024)
por: Dorn, Rebecca, et al.
Publicado: (2024)
Dialectal Coverage And Generalization in Arabic Speech Recognition
por: Djanibekov, Amirbek, et al.
Publicado: (2024)
por: Djanibekov, Amirbek, et al.
Publicado: (2024)
Doing More with Less: Data Augmentation for Sudanese Dialect Automatic Speech Recognition
por: Mansour, Ayman
Publicado: (2026)
por: Mansour, Ayman
Publicado: (2026)
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology
por: Sullivan, Peter, et al.
Publicado: (2026)
por: Sullivan, Peter, et al.
Publicado: (2026)
Streaming Speech-to-Text Translation with a SpeechLLM
por: Parcollet, Titouan, et al.
Publicado: (2026)
por: Parcollet, Titouan, et al.
Publicado: (2026)
Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
por: Oberkircher, Lena S., et al.
Publicado: (2026)
por: Oberkircher, Lena S., et al.
Publicado: (2026)
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
por: Zhao, Xingjian, et al.
Publicado: (2025)
por: Zhao, Xingjian, et al.
Publicado: (2025)
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
por: Shikhar, Sambal, et al.
Publicado: (2025)
por: Shikhar, Sambal, et al.
Publicado: (2025)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
por: Cornell, Samuele, et al.
Publicado: (2024)
por: Cornell, Samuele, et al.
Publicado: (2024)
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
por: Chen, Li-Wei, et al.
Publicado: (2024)
por: Chen, Li-Wei, et al.
Publicado: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
por: Liu, Henglyu, et al.
Publicado: (2025)
por: Liu, Henglyu, et al.
Publicado: (2025)
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data
por: Lajčinová, Bibiána, et al.
Publicado: (2024)
por: Lajčinová, Bibiána, et al.
Publicado: (2024)
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
por: Naouara, Hedi, et al.
Publicado: (2025)
por: Naouara, Hedi, et al.
Publicado: (2025)
Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text
por: Ruby, Ahmed, et al.
Publicado: (2026)
por: Ruby, Ahmed, et al.
Publicado: (2026)
SpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to Accelerate Your Speech-LLM Development
por: Wang, Minghan, et al.
Publicado: (2025)
por: Wang, Minghan, et al.
Publicado: (2025)
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects
por: Yang, Sicheng, et al.
Publicado: (2026)
por: Yang, Sicheng, et al.
Publicado: (2026)
Ejemplares similares
-
GATech at AbjadGenEval Shared Task: Multilingual Embeddings for Arabic Machine-Generated Text Classification
por: Khamis, Ahmed Khaled
Publicado: (2026) -
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
por: Dua, Karan, et al.
Publicado: (2025) -
GATech at AbjadMed: Bidirectional Encoders vs. Causal Decoders: Insights from 82-Class Arabic Medical Classification
por: Khamis, Ahmed Khaled
Publicado: (2026) -
Phonetic Modeling of Dialectal Variation in Vietnamese Speech
por: Hoang, Quan Ngoc, et al.
Publicado: (2026) -
Towards Zero-Shot Text-To-Speech for Arabic Dialects
por: Doan, Khai Duy, et al.
Publicado: (2024)