Hello Afrika: Speech Commands in Kinyarwanda
Fuente:
arXiv
Salvato in:
| Autori principali: | Igwegbe, George, Awojide, Martins, Bless, Mboh, Kadzo, Nirel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
di: Lachenani, Sidahmed, et al.
Pubblicazione: (2025)
di: Lachenani, Sidahmed, et al.
Pubblicazione: (2025)
Hello-Chat: Towards Realistic Social Audio Interactions
di: Hou, Yueran, et al.
Pubblicazione: (2026)
di: Hou, Yueran, et al.
Pubblicazione: (2026)
Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
di: Izotov, Yuriy, et al.
Pubblicazione: (2025)
di: Izotov, Yuriy, et al.
Pubblicazione: (2025)
Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
di: Quintas, Sebastião, et al.
Pubblicazione: (2024)
di: Quintas, Sebastião, et al.
Pubblicazione: (2024)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
di: Lin, Zijian, et al.
Pubblicazione: (2025)
di: Lin, Zijian, et al.
Pubblicazione: (2025)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
di: Kim, Jaeyoung, et al.
Pubblicazione: (2024)
di: Kim, Jaeyoung, et al.
Pubblicazione: (2024)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
di: Wang, Yongqi, et al.
Pubblicazione: (2023)
di: Wang, Yongqi, et al.
Pubblicazione: (2023)
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
di: Mogridge, Rhiannon, et al.
Pubblicazione: (2024)
di: Mogridge, Rhiannon, et al.
Pubblicazione: (2024)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
di: Wang, Helin, et al.
Pubblicazione: (2025)
di: Wang, Helin, et al.
Pubblicazione: (2025)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
di: Ji, Zhoulin, et al.
Pubblicazione: (2024)
di: Ji, Zhoulin, et al.
Pubblicazione: (2024)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
di: Lau, Hok-Shing, et al.
Pubblicazione: (2024)
di: Lau, Hok-Shing, et al.
Pubblicazione: (2024)
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
di: Bian, Weizhen, et al.
Pubblicazione: (2024)
di: Bian, Weizhen, et al.
Pubblicazione: (2024)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
di: Moell, Birger, et al.
Pubblicazione: (2025)
di: Moell, Birger, et al.
Pubblicazione: (2025)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
di: Shah, Neil, et al.
Pubblicazione: (2024)
di: Shah, Neil, et al.
Pubblicazione: (2024)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
di: Choi, Yerin, et al.
Pubblicazione: (2024)
di: Choi, Yerin, et al.
Pubblicazione: (2024)
MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection
di: Songyi, Li, et al.
Pubblicazione: (2026)
di: Songyi, Li, et al.
Pubblicazione: (2026)
TinyML for Speech Recognition
di: Barovic, Andrew, et al.
Pubblicazione: (2025)
di: Barovic, Andrew, et al.
Pubblicazione: (2025)
MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
di: Shah, Neil, et al.
Pubblicazione: (2024)
di: Shah, Neil, et al.
Pubblicazione: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
di: Shi, Hao, et al.
Pubblicazione: (2024)
di: Shi, Hao, et al.
Pubblicazione: (2024)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
di: Wang, Helin, et al.
Pubblicazione: (2025)
di: Wang, Helin, et al.
Pubblicazione: (2025)
Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People
di: Zhou, Haoshuai, et al.
Pubblicazione: (2025)
di: Zhou, Haoshuai, et al.
Pubblicazione: (2025)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
di: Wang, Xinsheng, et al.
Pubblicazione: (2025)
di: Wang, Xinsheng, et al.
Pubblicazione: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
DASB - Discrete Audio and Speech Benchmark
di: Mousavi, Pooneh, et al.
Pubblicazione: (2024)
di: Mousavi, Pooneh, et al.
Pubblicazione: (2024)
An Investigation of Incorporating Mamba for Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2024)
di: Chao, Rong, et al.
Pubblicazione: (2024)
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
di: Lemerle, Théodor, et al.
Pubblicazione: (2024)
di: Lemerle, Théodor, et al.
Pubblicazione: (2024)
Adaptive Duration Model for Text Speech Alignment
di: Cao, Junjie
Pubblicazione: (2025)
di: Cao, Junjie
Pubblicazione: (2025)
MathReader : Text-to-Speech for Mathematical Documents
di: Hyeon, Sieun, et al.
Pubblicazione: (2025)
di: Hyeon, Sieun, et al.
Pubblicazione: (2025)
Decoding Order Matters in Autoregressive Speech Synthesis
di: Zhao, Minghui, et al.
Pubblicazione: (2026)
di: Zhao, Minghui, et al.
Pubblicazione: (2026)
Study of the Performance of CEEMDAN in Underdetermined Speech Separation
di: Melhem, Rawad, et al.
Pubblicazione: (2024)
di: Melhem, Rawad, et al.
Pubblicazione: (2024)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
di: Ahn, Hyebin, et al.
Pubblicazione: (2025)
di: Ahn, Hyebin, et al.
Pubblicazione: (2025)
Adaptive Knowledge Distillation for Device-Directed Speech Detection
di: Chi, Hyung Gun, et al.
Pubblicazione: (2025)
di: Chi, Hyung Gun, et al.
Pubblicazione: (2025)
Temporal-Aware Iterative Speech Model for Dementia Detection
di: Ugwu, Chukwuemeka, et al.
Pubblicazione: (2025)
di: Ugwu, Chukwuemeka, et al.
Pubblicazione: (2025)
CIS-BWE: Chaos-Informed Speech Bandwidth Extension
di: Tamiti, Tarikul Islam, et al.
Pubblicazione: (2025)
di: Tamiti, Tarikul Islam, et al.
Pubblicazione: (2025)
Unlocking Speech Instruction Data Potential with Query Rewriting
di: Hei, Yonghua, et al.
Pubblicazione: (2025)
di: Hei, Yonghua, et al.
Pubblicazione: (2025)
Unsupervised Speech Enhancement using Data-defined Priors
di: Klement, Dominik, et al.
Pubblicazione: (2025)
di: Klement, Dominik, et al.
Pubblicazione: (2025)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
di: Vallés-Pérez, Ivan, et al.
Pubblicazione: (2023)
di: Vallés-Pérez, Ivan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
di: Lachenani, Sidahmed, et al.
Pubblicazione: (2025) -
Hello-Chat: Towards Realistic Social Audio Interactions
di: Hou, Yueran, et al.
Pubblicazione: (2026) -
Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
di: Izotov, Yuriy, et al.
Pubblicazione: (2025) -
Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
di: Quintas, Sebastião, et al.
Pubblicazione: (2024) -
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
di: Lin, Zijian, et al.
Pubblicazione: (2025)