PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Halpern, Bence Mark, Tienkamp, Thomas, Abur, Defne, Toda, Tomoki |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Quantifying the effect of speech pathology on automatic and human speaker verification
par: Halpern, Bence Mark, et autres
Publié: (2024)
par: Halpern, Bence Mark, et autres
Publié: (2024)
Towards explainable reference-free speech intelligibility evaluation of people with pathological speech
par: Halpern, Bence Mark, et autres
Publié: (2026)
par: Halpern, Bence Mark, et autres
Publié: (2026)
BlasBench: An Open Benchmark for Irish Speech Recognition
par: Raj, Jyoutir, et autres
Publié: (2026)
par: Raj, Jyoutir, et autres
Publié: (2026)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
par: Chowdhury, MD. Sagor, et autres
Publié: (2026)
par: Chowdhury, MD. Sagor, et autres
Publié: (2026)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
par: Ahn, Taekyung, et autres
Publié: (2024)
par: Ahn, Taekyung, et autres
Publié: (2024)
Measuring the Accuracy of Automatic Speech Recognition Solutions
par: Kuhn, Korbinian, et autres
Publié: (2024)
par: Kuhn, Korbinian, et autres
Publié: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
par: Cheng, Zhuangfei, et autres
Publié: (2025)
par: Cheng, Zhuangfei, et autres
Publié: (2025)
WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
par: Chowdhury, Aurchi, et autres
Publié: (2026)
par: Chowdhury, Aurchi, et autres
Publié: (2026)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
Communication Access Real-Time Translation Through Collaborative Correction of Automatic Speech Recognition
par: Kuhn, Korbinian, et autres
Publié: (2025)
par: Kuhn, Korbinian, et autres
Publié: (2025)
Everyday Speech in the Indian Subcontinent
par: P, Utkarsh
Publié: (2024)
par: P, Utkarsh
Publié: (2024)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
par: Feng, Ruitao, et autres
Publié: (2025)
par: Feng, Ruitao, et autres
Publié: (2025)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
par: Salvi, Giampiero
Publié: (2024)
par: Salvi, Giampiero
Publié: (2024)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
par: Adelson, Trevor, et autres
Publié: (2026)
par: Adelson, Trevor, et autres
Publié: (2026)
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
par: Christop, Iwona, et autres
Publié: (2026)
par: Christop, Iwona, et autres
Publié: (2026)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
par: Fang, Qingkai, et autres
Publié: (2024)
par: Fang, Qingkai, et autres
Publié: (2024)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
par: Robertson, Sean, et autres
Publié: (2023)
par: Robertson, Sean, et autres
Publié: (2023)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
par: Fang, Qingkai, et autres
Publié: (2024)
par: Fang, Qingkai, et autres
Publié: (2024)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
par: Ismail, Saifelden M.
Publié: (2025)
par: Ismail, Saifelden M.
Publié: (2025)
Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
par: Wang, Haoyu, et autres
Publié: (2025)
par: Wang, Haoyu, et autres
Publié: (2025)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
par: Jain, Sarthak, et autres
Publié: (2024)
par: Jain, Sarthak, et autres
Publié: (2024)
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
par: Chien, Sheng-You, et autres
Publié: (2026)
par: Chien, Sheng-You, et autres
Publié: (2026)
Diffuse or Confuse: A Diffusion Deepfake Speech Dataset
par: Firc, Anton, et autres
Publié: (2024)
par: Firc, Anton, et autres
Publié: (2024)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
par: Sharma, Manali, et autres
Publié: (2026)
par: Sharma, Manali, et autres
Publié: (2026)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
par: Wang, Shaowen, et autres
Publié: (2025)
par: Wang, Shaowen, et autres
Publié: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
par: Alamr, Meshal, et autres
Publié: (2026)
par: Alamr, Meshal, et autres
Publié: (2026)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
par: Wang, Hsuan-Yu, et autres
Publié: (2025)
par: Wang, Hsuan-Yu, et autres
Publié: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
par: Viswanathan, Janaki, et autres
Publié: (2025)
par: Viswanathan, Janaki, et autres
Publié: (2025)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
par: Li, Pengcheng, et autres
Publié: (2024)
par: Li, Pengcheng, et autres
Publié: (2024)
Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset
par: Marie, Ambre, et autres
Publié: (2025)
par: Marie, Ambre, et autres
Publié: (2025)
A Sociolinguistic Analysis of Automatic Speech Recognition Bias in Newcastle English
par: Serditova, Dana, et autres
Publié: (2026)
par: Serditova, Dana, et autres
Publié: (2026)
Syllable based DNN-HMM Cantonese Speech to Text System
par: Wong, Timothy, et autres
Publié: (2024)
par: Wong, Timothy, et autres
Publié: (2024)
VoiceSHIELD-Small: Real-Time Malicious Speech Detection and Transcription
par: Ranjan, Sumit, et autres
Publié: (2026)
par: Ranjan, Sumit, et autres
Publié: (2026)
Improving French Synthetic Speech Quality via SSML Prosody Control
par: Ouali, Nassima Ould, et autres
Publié: (2025)
par: Ouali, Nassima Ould, et autres
Publié: (2025)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
par: Fang, Qingkai, et autres
Publié: (2024)
par: Fang, Qingkai, et autres
Publié: (2024)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
par: Simões, Lucca Emmanuel Pineli, et autres
Publié: (2024)
par: Simões, Lucca Emmanuel Pineli, et autres
Publié: (2024)
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit
par: Soni, Aniket Abhishek
Publié: (2025)
par: Soni, Aniket Abhishek
Publié: (2025)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
par: Bhadra, Dipayan, et autres
Publié: (2025)
par: Bhadra, Dipayan, et autres
Publié: (2025)
Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer
par: Halpern, Bence Mark, et autres
Publié: (2025)
par: Halpern, Bence Mark, et autres
Publié: (2025)
Documents similaires
-
Quantifying the effect of speech pathology on automatic and human speaker verification
par: Halpern, Bence Mark, et autres
Publié: (2024) -
Towards explainable reference-free speech intelligibility evaluation of people with pathological speech
par: Halpern, Bence Mark, et autres
Publié: (2026) -
BlasBench: An Open Benchmark for Irish Speech Recognition
par: Raj, Jyoutir, et autres
Publié: (2026) -
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
par: Chowdhury, MD. Sagor, et autres
Publié: (2026) -
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
par: Ahn, Taekyung, et autres
Publié: (2024)