SeMaScore : a new evaluation metric for automatic speech recognition tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Sasindran, Zitha, Yelchuri, Harsha, Prabhakar, T. V. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Zipformer: A faster and better encoder for automatic speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2023)
di: Yao, Zengwei, et al.
Pubblicazione: (2023)
Robustifying automatic speech recognition by extracting slowly varying features
di: Pizarro, Matías, et al.
Pubblicazione: (2021)
di: Pizarro, Matías, et al.
Pubblicazione: (2021)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
di: Leglaive, Simon, et al.
Pubblicazione: (2024)
di: Leglaive, Simon, et al.
Pubblicazione: (2024)
CR-CTC: Consistency regularization on CTC for improved speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
Late fusion ensembles for speech recognition on diverse input audio representations
di: Jezidžić, Marin, et al.
Pubblicazione: (2024)
di: Jezidžić, Marin, et al.
Pubblicazione: (2024)
Introduction to speech recognition
di: Dauphin, Gabriel
Pubblicazione: (2024)
di: Dauphin, Gabriel
Pubblicazione: (2024)
An Attention Long Short-Term Memory based system for automatic classification of speech intelligibility
di: Fernández-Díaz, Miguel, et al.
Pubblicazione: (2024)
di: Fernández-Díaz, Miguel, et al.
Pubblicazione: (2024)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
di: Pepino, Leonardo, et al.
Pubblicazione: (2024)
di: Pepino, Leonardo, et al.
Pubblicazione: (2024)
CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech
di: Pahar, Madhurananda, et al.
Pubblicazione: (2025)
di: Pahar, Madhurananda, et al.
Pubblicazione: (2025)
Acoustic characterization of speech rhythm: going beyond metrics with recurrent neural networks
di: Deloche, François, et al.
Pubblicazione: (2024)
di: Deloche, François, et al.
Pubblicazione: (2024)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
di: Sadok, Samir, et al.
Pubblicazione: (2023)
di: Sadok, Samir, et al.
Pubblicazione: (2023)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
di: Zhang, Yiru, et al.
Pubblicazione: (2025)
di: Zhang, Yiru, et al.
Pubblicazione: (2025)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
di: Kim, Yunsik, et al.
Pubblicazione: (2025)
di: Kim, Yunsik, et al.
Pubblicazione: (2025)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
di: Gong, Rong, et al.
Pubblicazione: (2024)
di: Gong, Rong, et al.
Pubblicazione: (2024)
Selfsupervised learning for pathological speech detection
di: Sheikh, Shakeel Ahmad
Pubblicazione: (2024)
di: Sheikh, Shakeel Ahmad
Pubblicazione: (2024)
Towards the Synthesis of Non-speech Vocalizations
di: Hoq, Enjamamul, et al.
Pubblicazione: (2024)
di: Hoq, Enjamamul, et al.
Pubblicazione: (2024)
Cough activity detection for automatic tuberculosis screening
di: van Vüren, Joshua Jansen, et al.
Pubblicazione: (2026)
di: van Vüren, Joshua Jansen, et al.
Pubblicazione: (2026)
Optimising MFCC parameters for the automatic detection of respiratory diseases
di: Yan, Yuyang, et al.
Pubblicazione: (2024)
di: Yan, Yuyang, et al.
Pubblicazione: (2024)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
di: Chen, Szu-Jui, et al.
Pubblicazione: (2026)
di: Chen, Szu-Jui, et al.
Pubblicazione: (2026)
Audio-based automatic mating success prediction of giant pandas
di: Yan, WeiRan, et al.
Pubblicazione: (2019)
di: Yan, WeiRan, et al.
Pubblicazione: (2019)
Single-channel speech enhancement using learnable loss mixup
di: Chang, Oscar, et al.
Pubblicazione: (2023)
di: Chang, Oscar, et al.
Pubblicazione: (2023)
Benchmarks and leaderboards for sound demixing tasks
di: Solovyev, Roman, et al.
Pubblicazione: (2023)
di: Solovyev, Roman, et al.
Pubblicazione: (2023)
Distilling a speech and music encoder with task arithmetic
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
Boosting keyword spotting through on-device learnable user speech characteristics
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
Generalizable speech deepfake detection via meta-learned LoRA
di: Laakkonen, Janne, et al.
Pubblicazione: (2025)
di: Laakkonen, Janne, et al.
Pubblicazione: (2025)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
di: Araiza-Illan, Gloria, et al.
Pubblicazione: (2023)
di: Araiza-Illan, Gloria, et al.
Pubblicazione: (2023)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
Context-aware child-directed speech detection from long-form recordings
di: Charlot, Théo, et al.
Pubblicazione: (2026)
di: Charlot, Théo, et al.
Pubblicazione: (2026)
Dementia classification from spontaneous speech using wrapper-based feature selection
di: Niemelä, Marko, et al.
Pubblicazione: (2025)
di: Niemelä, Marko, et al.
Pubblicazione: (2025)
Analyzing and reducing the synthetic-to-real transfer gap in Music Information Retrieval: the task of automatic drum transcription
di: Zehren, Mickaël, et al.
Pubblicazione: (2024)
di: Zehren, Mickaël, et al.
Pubblicazione: (2024)
Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
di: Maisonneuve, Malo, et al.
Pubblicazione: (2024)
di: Maisonneuve, Malo, et al.
Pubblicazione: (2024)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2023)
di: Maharana, Sarthak Kumar, et al.
Pubblicazione: (2023)
Test-Time Training for Depression Detection
di: Dumpala, Sri Harsha, et al.
Pubblicazione: (2024)
di: Dumpala, Sri Harsha, et al.
Pubblicazione: (2024)
Adaptive ship-radiated noise recognition with learnable fine-grained wavelet transform
di: Xie, Yuan, et al.
Pubblicazione: (2023)
di: Xie, Yuan, et al.
Pubblicazione: (2023)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
di: Wright, George August, et al.
Pubblicazione: (2023)
di: Wright, George August, et al.
Pubblicazione: (2023)
Determining the severity of Parkinson's disease in patients using a multi task neural network
di: García-Ordás, María Teresa, et al.
Pubblicazione: (2024)
di: García-Ordás, María Teresa, et al.
Pubblicazione: (2024)
HRTF Estimation using a Score-based Prior
di: Thuillier, Etienne, et al.
Pubblicazione: (2024)
di: Thuillier, Etienne, et al.
Pubblicazione: (2024)
KinSPEAK: Improving speech recognition for Kinyarwanda via semi-supervised learning methods
di: Nzeyimana, Antoine
Pubblicazione: (2023)
di: Nzeyimana, Antoine
Pubblicazione: (2023)
Documenti analoghi
-
Zipformer: A faster and better encoder for automatic speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2023) -
Robustifying automatic speech recognition by extracting slowly varying features
di: Pizarro, Matías, et al.
Pubblicazione: (2021) -
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
di: Maiti, Soumi, et al.
Pubblicazione: (2023) -
Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
di: Leglaive, Simon, et al.
Pubblicazione: (2024) -
CR-CTC: Consistency regularization on CTC for improved speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2024)