Arabic Little STT: Arabic Children Speech Recognition Dataset
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Alkadri, Mouhand, Desouki, Dania, Jallad, Khloud Al |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
von: Darwish, Hazem, et al.
Veröffentlicht: (2024)
von: Darwish, Hazem, et al.
Veröffentlicht: (2024)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
ArEEG_Chars: Dataset for Envisioned Speech Recognition using EEG for Arabic Characters
von: Darwish, Hazem, et al.
Veröffentlicht: (2024)
von: Darwish, Hazem, et al.
Veröffentlicht: (2024)
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
von: Khalil, Mohammad Amer, et al.
Veröffentlicht: (2026)
von: Khalil, Mohammad Amer, et al.
Veröffentlicht: (2026)
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
von: Jallad, Khloud AL, et al.
Veröffentlicht: (2025)
von: Jallad, Khloud AL, et al.
Veröffentlicht: (2025)
Morse Code-Enabled Speech Recognition for Individuals with Visual and Hearing Impairments
von: Choudhury, Ritabrata Roy
Veröffentlicht: (2024)
von: Choudhury, Ritabrata Roy
Veröffentlicht: (2024)
What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
von: Ibrahim, Adham, et al.
Veröffentlicht: (2024)
von: Ibrahim, Adham, et al.
Veröffentlicht: (2024)
AIN: The Arabic INclusive Large Multimodal Model
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
Voting-based Multimodal Automatic Deception Detection
von: Touma, Lana, et al.
Veröffentlicht: (2023)
von: Touma, Lana, et al.
Veröffentlicht: (2023)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
Do AI Voices Learn Social Nuances? A Case of Politeness and Speech Rate
von: Rabin, Eyal, et al.
Veröffentlicht: (2025)
von: Rabin, Eyal, et al.
Veröffentlicht: (2025)
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
von: Liu, Tianyun
Veröffentlicht: (2025)
von: Liu, Tianyun
Veröffentlicht: (2025)
Call2Instruct: Automated Pipeline for Generating Q&A Datasets from Call Center Recordings for LLM Fine-Tuning
von: Echeverria, Alex, et al.
Veröffentlicht: (2025)
von: Echeverria, Alex, et al.
Veröffentlicht: (2025)
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
Adaptation and Optimization of Automatic Speech Recognition (ASR) for the Maritime Domain in the Field of VHF Communication
von: Nakilcioglu, Emin Cagatay, et al.
Veröffentlicht: (2023)
von: Nakilcioglu, Emin Cagatay, et al.
Veröffentlicht: (2023)
Spontaneous Informal Speech Dataset for Punctuation Restoration
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks
von: Daoud, Mouath Abu, et al.
Veröffentlicht: (2025)
von: Daoud, Mouath Abu, et al.
Veröffentlicht: (2025)
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
Evaluating Human-AI Interaction via Usability, User Experience and Acceptance Measures for MMM-C: A Creative AI System for Music Composition
von: Tchemeube, Renaud Bougueng, et al.
Veröffentlicht: (2025)
von: Tchemeube, Renaud Bougueng, et al.
Veröffentlicht: (2025)
Voice "Cloning" is Style Transfer
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2026)
Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection
von: Chandra, Joydeep
Veröffentlicht: (2026)
von: Chandra, Joydeep
Veröffentlicht: (2026)
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
von: Gheffari, Youcef Soufiane, et al.
Veröffentlicht: (2026)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
von: Dietrich, Juergen
Veröffentlicht: (2026)
von: Dietrich, Juergen
Veröffentlicht: (2026)
EmoHeal: An End-to-End System for Personalized Therapeutic Music Retrieval from Fine-grained Emotions
von: Wan, Xinchen, et al.
Veröffentlicht: (2025)
von: Wan, Xinchen, et al.
Veröffentlicht: (2025)
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
von: Xie, Zhifei, et al.
Veröffentlicht: (2024)
von: Xie, Zhifei, et al.
Veröffentlicht: (2024)
EmoAugNet: A Signal-Augmented Hybrid CNN-LSTM Framework for Speech Emotion Recognition
von: Paul, Durjoy Chandra, et al.
Veröffentlicht: (2025)
von: Paul, Durjoy Chandra, et al.
Veröffentlicht: (2025)
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
PRECISE Framework: GPT-based Text For Improved Readability, Reliability, and Understandability of Radiology Reports For Patient-Centered Care
von: Tripathi, Satvik, et al.
Veröffentlicht: (2024)
von: Tripathi, Satvik, et al.
Veröffentlicht: (2024)
Deep Learning Models in Speech Recognition: Measuring GPU Energy Consumption, Impact of Noise and Model Quantization for Edge Deployment
von: Chakravarty, Aditya
Veröffentlicht: (2024)
von: Chakravarty, Aditya
Veröffentlicht: (2024)
AraModernBERT: Transtokenized Initialization and Long-Context Encoder Modeling for Arabic
von: Elshehy, Omar, et al.
Veröffentlicht: (2026)
von: Elshehy, Omar, et al.
Veröffentlicht: (2026)
Introducing MeMo: A Multimodal Dataset for Memory Modelling in Multiparty Conversations
von: Tsfasman, Maria, et al.
Veröffentlicht: (2024)
von: Tsfasman, Maria, et al.
Veröffentlicht: (2024)
LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts
von: Mohammadi, Seyedali, et al.
Veröffentlicht: (2025)
von: Mohammadi, Seyedali, et al.
Veröffentlicht: (2025)
Acoustic and perceptual differences between standard and accented speech and their voice clones
von: Yang, Tianle, et al.
Veröffentlicht: (2026)
von: Yang, Tianle, et al.
Veröffentlicht: (2026)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
von: Benster, Tyler, et al.
Veröffentlicht: (2024)
von: Benster, Tyler, et al.
Veröffentlicht: (2024)
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
von: Chen, Maximillian, et al.
Veröffentlicht: (2026)
von: Chen, Maximillian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
von: Darwish, Hazem, et al.
Veröffentlicht: (2024) -
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026) -
ArEEG_Chars: Dataset for Envisioned Speech Recognition using EEG for Arabic Characters
von: Darwish, Hazem, et al.
Veröffentlicht: (2024) -
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
von: Khalil, Mohammad Amer, et al.
Veröffentlicht: (2026) -
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
von: Jallad, Khloud AL, et al.
Veröffentlicht: (2025)