Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
Fuente:
arXiv
Salvato in:
| Autori principali: | Azad, Asif, Shanto, MD Sadik Hossain, Hossain, Mohammad Sadat, Alwuqaysi, Bdour, Boughorbel, Sabri, Bokhari, Yahya, Aljouie, Abdulrhman, Sindi, Ayah Othman, Hoque, Ehsan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Design and Evaluation of a Culturally Adapted Multimodal Virtual Agent for PTSD Screening
di: Ozel, Cengiz, et al.
Pubblicazione: (2026)
di: Ozel, Cengiz, et al.
Pubblicazione: (2026)
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
di: Azad, Asif, et al.
Pubblicazione: (2025)
di: Azad, Asif, et al.
Pubblicazione: (2025)
AttMetNet: Attention-Enhanced Deep Neural Network for Methane Plume Detection in Sentinel-2 Satellite Imagery
di: Ahsan, Rakib, et al.
Pubblicazione: (2025)
di: Ahsan, Rakib, et al.
Pubblicazione: (2025)
A Novel Fusion Architecture for PD Detection Using Semi-Supervised Speech Embeddings
di: Adnan, Tariq, et al.
Pubblicazione: (2024)
di: Adnan, Tariq, et al.
Pubblicazione: (2024)
HyMNet: a Multimodal Deep Learning System for Hypertension Classification using Fundus Photographs and Cardiometabolic Risk Factors
di: Baharoon, Mohammed, et al.
Pubblicazione: (2023)
di: Baharoon, Mohammed, et al.
Pubblicazione: (2023)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
di: Rashid, Md Rafi Ur, et al.
Pubblicazione: (2026)
di: Rashid, Md Rafi Ur, et al.
Pubblicazione: (2026)
Improving Language Models Trained on Translated Data with Continual Pre-Training and Dictionary Learning Analysis
di: Boughorbel, Sabri, et al.
Pubblicazione: (2024)
di: Boughorbel, Sabri, et al.
Pubblicazione: (2024)
AraS2P: Arabic Speech-to-Phonemes System
di: Matar, Bassam, et al.
Pubblicazione: (2025)
di: Matar, Bassam, et al.
Pubblicazione: (2025)
Cochleagram-based Noise Adapted Speaker Identification System for Distorted Speech
di: Ahmed, Sabbir, et al.
Pubblicazione: (2025)
di: Ahmed, Sabbir, et al.
Pubblicazione: (2025)
Gradient Masters at BLP-2025 Task 1: Advancing Low-Resource NLP for Bengali using Ensemble-Based Adversarial Training for Hate Speech Detection
di: Hoque, Syed Mohaiminul, et al.
Pubblicazione: (2025)
di: Hoque, Syed Mohaiminul, et al.
Pubblicazione: (2025)
DFCon: Attention-Driven Supervised Contrastive Learning for Robust Deepfake Detection
di: Shanto, MD Sadik Hossain, et al.
Pubblicazione: (2025)
di: Shanto, MD Sadik Hossain, et al.
Pubblicazione: (2025)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
di: Hossain, Eftekhar, et al.
Pubblicazione: (2024)
di: Hossain, Eftekhar, et al.
Pubblicazione: (2024)
Evaluating General Purpose Vision Foundation Models for Medical Image Analysis: An Experimental Study of DINOv2 on Radiology Benchmarks
di: Baharoon, Mohammed, et al.
Pubblicazione: (2023)
di: Baharoon, Mohammed, et al.
Pubblicazione: (2023)
ViSpeechFormer: A Phonemic Approach for Vietnamese Automatic Speech Recognition
di: Nguyen, Khoa Anh, et al.
Pubblicazione: (2026)
di: Nguyen, Khoa Anh, et al.
Pubblicazione: (2026)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
di: Koriyama, Tomoki
Pubblicazione: (2025)
di: Koriyama, Tomoki
Pubblicazione: (2025)
CARMA: Comprehensive Automatically-annotated Reddit Mental Health Dataset for Arabic
di: Mankarious, Saad, et al.
Pubblicazione: (2025)
di: Mankarious, Saad, et al.
Pubblicazione: (2025)
Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing
di: Boughorbel, Sabri, et al.
Pubblicazione: (2025)
di: Boughorbel, Sabri, et al.
Pubblicazione: (2025)
A comparative study on the effect of commercial fish feeds on the growth of Thai pangas, Pangasius hypophthalmus
di: Kader, M.A., et al.
Pubblicazione: (2003)
di: Kader, M.A., et al.
Pubblicazione: (2003)
A Phoneme-Scale Assessment of Multichannel Speech Enhancement Algorithms
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2024)
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2024)
Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
di: Salvi, Davide, et al.
Pubblicazione: (2025)
di: Salvi, Davide, et al.
Pubblicazione: (2025)
Real-Time Detection and Analysis of Vehicles and Pedestrians using Deep Learning
di: Sadik, Md Nahid, et al.
Pubblicazione: (2024)
di: Sadik, Md Nahid, et al.
Pubblicazione: (2024)
SpeechAlign: Aligning Speech Generation to Human Preferences
di: Zhang, Dong, et al.
Pubblicazione: (2024)
di: Zhang, Dong, et al.
Pubblicazione: (2024)
Advancing Assistive Robotics: Multi-Modal Navigation and Biophysical Monitoring for Next-Generation Wheelchairs
di: Hossain, Md. Anowar, et al.
Pubblicazione: (2026)
di: Hossain, Md. Anowar, et al.
Pubblicazione: (2026)
Arabic Little STT: Arabic Children Speech Recognition Dataset
di: Alkadri, Mouhand, et al.
Pubblicazione: (2025)
di: Alkadri, Mouhand, et al.
Pubblicazione: (2025)
Multitask Learning for Grapheme-to-Phoneme Conversion of Anglicisms in German Speech Recognition
di: Pritzen, Julia, et al.
Pubblicazione: (2021)
di: Pritzen, Julia, et al.
Pubblicazione: (2021)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
di: Karapiperis, Sotirios, et al.
Pubblicazione: (2024)
di: Karapiperis, Sotirios, et al.
Pubblicazione: (2024)
MEGConformer: Conformer-Based MEG Decoder for Robust Speech and Phoneme Classification
di: de Zuazo, Xabier, et al.
Pubblicazione: (2025)
di: de Zuazo, Xabier, et al.
Pubblicazione: (2025)
CUPE: Contextless Universal Phoneme Encoder for Language-Agnostic Speech Processing
di: Rehman, Abdul, et al.
Pubblicazione: (2025)
di: Rehman, Abdul, et al.
Pubblicazione: (2025)
Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2025)
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2025)
Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech
di: Kolani, Yakov, et al.
Pubblicazione: (2025)
di: Kolani, Yakov, et al.
Pubblicazione: (2025)
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
di: Xue, Jun, et al.
Pubblicazione: (2026)
di: Xue, Jun, et al.
Pubblicazione: (2026)
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
di: Biswas, Subrata, et al.
Pubblicazione: (2025)
di: Biswas, Subrata, et al.
Pubblicazione: (2025)
Memory Under Siege: A Comprehensive Survey of Side-Channel Attacks on Memory
di: Hassan, MD Mahady, et al.
Pubblicazione: (2025)
di: Hassan, MD Mahady, et al.
Pubblicazione: (2025)
Demographic-Aware Transfer Learning for Sleep Stage Classification in Clinical Polysomnography
di: Hossain, S M Asif, et al.
Pubblicazione: (2026)
di: Hossain, S M Asif, et al.
Pubblicazione: (2026)
InfiltrNet: Dual-Branch CNN-Transformer Architecture for Brain Tumor Infiltration Risk Prediction
di: Hossain, S M Asif, et al.
Pubblicazione: (2026)
di: Hossain, S M Asif, et al.
Pubblicazione: (2026)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
di: Chowdhury, MD. Sagor, et al.
Pubblicazione: (2026)
di: Chowdhury, MD. Sagor, et al.
Pubblicazione: (2026)
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
di: Zahid, Anwar Hossain, et al.
Pubblicazione: (2025)
di: Zahid, Anwar Hossain, et al.
Pubblicazione: (2025)
EmoHopeSpeech: An Annotated Dataset of Emotions and Hope Speech in English and Arabic
di: Zaghouani, Wajdi, et al.
Pubblicazione: (2025)
di: Zaghouani, Wajdi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Design and Evaluation of a Culturally Adapted Multimodal Virtual Agent for PTSD Screening
di: Ozel, Cengiz, et al.
Pubblicazione: (2026) -
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
di: Azad, Asif, et al.
Pubblicazione: (2025) -
AttMetNet: Attention-Enhanced Deep Neural Network for Methane Plume Detection in Sentinel-2 Satellite Imagery
di: Ahsan, Rakib, et al.
Pubblicazione: (2025) -
A Novel Fusion Architecture for PD Detection Using Semi-Supervised Speech Embeddings
di: Adnan, Tariq, et al.
Pubblicazione: (2024) -
HyMNet: a Multimodal Deep Learning System for Hypertension Classification using Fundus Photographs and Cardiometabolic Risk Factors
di: Baharoon, Mohammed, et al.
Pubblicazione: (2023)