Guardado en:
| Autores principales: | Kheir, Yassine El, Mubarak, Hamdy, Ali, Ahmed, Chowdhury, Shammur Absar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2408.02430 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Speech Representation Analysis based on Inter- and Intra-Model Similarities
por: Kheir, Yassine El, et al.
Publicado: (2024)
por: Kheir, Yassine El, et al.
Publicado: (2024)
IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA)
por: Kheir, Yassine El, et al.
Publicado: (2026)
por: Kheir, Yassine El, et al.
Publicado: (2026)
CAFE A Novel Code switching Dataset for Algerian Dialect French and English
por: Lachemat, Houssam Eddine-Othman, et al.
Publicado: (2024)
por: Lachemat, Houssam Eddine-Othman, et al.
Publicado: (2024)
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
por: Sukhadia, Vrunda N., et al.
Publicado: (2026)
por: Sukhadia, Vrunda N., et al.
Publicado: (2026)
Children's Speech Recognition through Discrete Token Enhancement
por: Sukhadia, Vrunda N., et al.
Publicado: (2024)
por: Sukhadia, Vrunda N., et al.
Publicado: (2024)
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs
por: Bhatti, Hunzalah Hassan, et al.
Publicado: (2026)
por: Bhatti, Hunzalah Hassan, et al.
Publicado: (2026)
BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention
por: Kheir, Yassine El, et al.
Publicado: (2025)
por: Kheir, Yassine El, et al.
Publicado: (2025)
Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study
por: Kheir, Yassine El, et al.
Publicado: (2025)
por: Kheir, Yassine El, et al.
Publicado: (2025)
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection
por: Kheir, Yassine El, et al.
Publicado: (2025)
por: Kheir, Yassine El, et al.
Publicado: (2025)
Automatic Assessment of Dysarthria Using Audio-visual Vowel Graph Attention Network
por: Liu, Xiaokang, et al.
Publicado: (2024)
por: Liu, Xiaokang, et al.
Publicado: (2024)
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
por: Naouara, Hedi, et al.
Publicado: (2025)
por: Naouara, Hedi, et al.
Publicado: (2025)
Dialectal Coverage And Generalization in Arabic Speech Recognition
por: Djanibekov, Amirbek, et al.
Publicado: (2024)
por: Djanibekov, Amirbek, et al.
Publicado: (2024)
Generalizable Audio Spoofing Detection using Non-Semantic Representations
por: Das, Arnab, et al.
Publicado: (2025)
por: Das, Arnab, et al.
Publicado: (2025)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
por: Doan, Khai Duy, et al.
Publicado: (2024)
por: Doan, Khai Duy, et al.
Publicado: (2024)
Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection
por: Kheir, Yassine El, et al.
Publicado: (2025)
por: Kheir, Yassine El, et al.
Publicado: (2025)
DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection
por: Kheir, Yassine El, et al.
Publicado: (2026)
por: Kheir, Yassine El, et al.
Publicado: (2026)
Automatic Speech Recognition with BERT and CTC Transformers: A Review
por: Djeffal, Noussaiba, et al.
Publicado: (2024)
por: Djeffal, Noussaiba, et al.
Publicado: (2024)
MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
por: Ali, Zien Sheikh, et al.
Publicado: (2026)
por: Ali, Zien Sheikh, et al.
Publicado: (2026)
Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
por: Abdullah, Badr M., et al.
Publicado: (2025)
por: Abdullah, Badr M., et al.
Publicado: (2025)
Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis
por: Chen, Yushen, et al.
Publicado: (2026)
por: Chen, Yushen, et al.
Publicado: (2026)
You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel Properties
por: Tuttösí, Paige, et al.
Publicado: (2025)
por: Tuttösí, Paige, et al.
Publicado: (2025)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
por: Özyilmaz, Ömer Tarik, et al.
Publicado: (2025)
por: Özyilmaz, Ömer Tarik, et al.
Publicado: (2025)
Hybrid Deep Learning and Signal Processing for Arabic Dialect Recognition in Low-Resource Settings
por: Al-Shwayyat, Ghazal, et al.
Publicado: (2025)
por: Al-Shwayyat, Ghazal, et al.
Publicado: (2025)
Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
por: Ding, Jiani, et al.
Publicado: (2025)
por: Ding, Jiani, et al.
Publicado: (2025)
A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
por: Blaschke, Verena, et al.
Publicado: (2025)
por: Blaschke, Verena, et al.
Publicado: (2025)
Dolphin-CN-Dialect: Where Chinese Dialects Matter
por: Meng, Yangyang, et al.
Publicado: (2026)
por: Meng, Yangyang, et al.
Publicado: (2026)
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
por: Kheddar, Hamza, et al.
Publicado: (2024)
por: Kheddar, Hamza, et al.
Publicado: (2024)
From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
por: Ersoy, Asım, et al.
Publicado: (2025)
por: Ersoy, Asım, et al.
Publicado: (2025)
LAMA-UT: Language Agnostic Multilingual ASR through Orthography Unification and Language-Specific Transliteration
por: Lee, Sangmin, et al.
Publicado: (2024)
por: Lee, Sangmin, et al.
Publicado: (2024)
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
por: Li, Baihan, et al.
Publicado: (2024)
por: Li, Baihan, et al.
Publicado: (2024)
Automatic Sound Event Detection and Classification of Great Ape Calls Using Neural Networks
por: Jiang, Zifan, et al.
Publicado: (2023)
por: Jiang, Zifan, et al.
Publicado: (2023)
Improving the Robustness and Clinical Applicability of Automatic Respiratory Sound Classification Using Deep Learning-Based Audio Enhancement: Algorithm Development and Validation
por: Tzeng, Jing-Tong, et al.
Publicado: (2024)
por: Tzeng, Jing-Tong, et al.
Publicado: (2024)
Arabic TTS with FastPitch: Reproducible Baselines, Adversarial Training, and Oversmoothing Analysis
por: Nippert, Lars
Publicado: (2025)
por: Nippert, Lars
Publicado: (2025)
Adaptive Representations of Sound for Automatic Insect Recognition
por: Faiß, Marius, et al.
Publicado: (2023)
por: Faiß, Marius, et al.
Publicado: (2023)
Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
por: Salman, Ali N., et al.
Publicado: (2024)
por: Salman, Ali N., et al.
Publicado: (2024)
Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
por: Gerazov, Branislav, et al.
Publicado: (2025)
por: Gerazov, Branislav, et al.
Publicado: (2025)
Classification of Short Segment Pediatric Heart Sounds Based on a Transformer-Based Convolutional Neural Network
por: Hassanuzzaman, Md, et al.
Publicado: (2024)
por: Hassanuzzaman, Md, et al.
Publicado: (2024)
A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR
por: Biswas, Swadhin, et al.
Publicado: (2025)
por: Biswas, Swadhin, et al.
Publicado: (2025)
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
por: Xie, Hanke, et al.
Publicado: (2025)
por: Xie, Hanke, et al.
Publicado: (2025)
Automatic Inspection Based on Switch Sounds of Electric Point Machines
por: Shibata, Ayano, et al.
Publicado: (2025)
por: Shibata, Ayano, et al.
Publicado: (2025)
Ejemplares similares
-
Speech Representation Analysis based on Inter- and Intra-Model Similarities
por: Kheir, Yassine El, et al.
Publicado: (2024) -
IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA)
por: Kheir, Yassine El, et al.
Publicado: (2026) -
CAFE A Novel Code switching Dataset for Algerian Dialect French and English
por: Lachemat, Houssam Eddine-Othman, et al.
Publicado: (2024) -
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
por: Sukhadia, Vrunda N., et al.
Publicado: (2026) -
Children's Speech Recognition through Discrete Token Enhancement
por: Sukhadia, Vrunda N., et al.
Publicado: (2024)