Evaluating ASR robustness to spontaneous speech errors: A study of WhisperX using a Speech Error Database
Fuente:
arXiv
Guardado en:
| Autores principales: | Alderete, John, Hui, Macarious Kin Fung, Mohan, Aanchan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
por: Hui, Macarious, et al.
Publicado: (2024)
por: Hui, Macarious, et al.
Publicado: (2024)
WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
por: Chowdhury, Aurchi, et al.
Publicado: (2026)
por: Chowdhury, Aurchi, et al.
Publicado: (2026)
A new approach for fine-tuning sentence transformers for intent classification and out-of-scope detection tasks
por: Zhang, Tianyi, et al.
Publicado: (2024)
por: Zhang, Tianyi, et al.
Publicado: (2024)
Human Latency Conversational Turns for Spoken Avatar Systems
por: Jacoby, Derek, et al.
Publicado: (2024)
por: Jacoby, Derek, et al.
Publicado: (2024)
Fine-tuning Whisper for Pashto ASR: strategies and scale
por: Rahman, Hanif
Publicado: (2026)
por: Rahman, Hanif
Publicado: (2026)
Configurable Multilingual ASR with Speech Summary Representations
por: Zhu, Harrison, et al.
Publicado: (2024)
por: Zhu, Harrison, et al.
Publicado: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
Extending Whisper with prompt tuning to target-speaker ASR
por: Ma, Hao, et al.
Publicado: (2023)
por: Ma, Hao, et al.
Publicado: (2023)
Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
por: Wang, Yingzhi, et al.
Publicado: (2025)
por: Wang, Yingzhi, et al.
Publicado: (2025)
Can Whisper perform speech-based in-context learning?
por: Wang, Siyin, et al.
Publicado: (2023)
por: Wang, Siyin, et al.
Publicado: (2023)
Whisper: Courtside Edition Enhancing ASR Performance Through LLM-Driven Context Generation
por: Ron, Yonathan, et al.
Publicado: (2026)
por: Ron, Yonathan, et al.
Publicado: (2026)
Careless Whisper: Speech-to-Text Hallucination Harms
por: Koenecke, Allison, et al.
Publicado: (2024)
por: Koenecke, Allison, et al.
Publicado: (2024)
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
por: Xu, Tianyi, et al.
Publicado: (2024)
por: Xu, Tianyi, et al.
Publicado: (2024)
On the Role of Encoder Depth: Pruning Whisper and LoRA Fine-Tuning in SLAM-ASR
por: Kolluri, Ganesh Pavan Kartikeya Bharadwaj, et al.
Publicado: (2026)
por: Kolluri, Ganesh Pavan Kartikeya Bharadwaj, et al.
Publicado: (2026)
LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
por: Song, Zheshu, et al.
Publicado: (2024)
por: Song, Zheshu, et al.
Publicado: (2024)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
por: Thorbecke, Iuliia, et al.
Publicado: (2024)
por: Thorbecke, Iuliia, et al.
Publicado: (2024)
Quantizing Whisper-small: How design choices affect ASR performance
por: Söhler, Arthur, et al.
Publicado: (2025)
por: Söhler, Arthur, et al.
Publicado: (2025)
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
por: Orhon, Atila, et al.
Publicado: (2025)
por: Orhon, Atila, et al.
Publicado: (2025)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
por: Wei, Victor Junqiu, et al.
Publicado: (2024)
por: Wei, Victor Junqiu, et al.
Publicado: (2024)
How much speech data is necessary for ASR in African languages? An evaluation of data scaling in Kinyarwanda and Kikuyu
por: Akera, Benjamin, et al.
Publicado: (2025)
por: Akera, Benjamin, et al.
Publicado: (2025)
Advocating Character Error Rate for Multilingual ASR Evaluation
por: K, Thennal D, et al.
Publicado: (2024)
por: K, Thennal D, et al.
Publicado: (2024)
From Speech to Subtitles: Evaluating ASR Models in Subtitling Italian Television Programs
por: Lucca, Alessandro, et al.
Publicado: (2025)
por: Lucca, Alessandro, et al.
Publicado: (2025)
Whisper-UT: A Unified Translation Framework for Speech and Text
por: Xiao, Cihan, et al.
Publicado: (2025)
por: Xiao, Cihan, et al.
Publicado: (2025)
ASR Error Correction using Large Language Models
por: Ma, Rao, et al.
Publicado: (2024)
por: Ma, Rao, et al.
Publicado: (2024)
PhoWhisper: Automatic Speech Recognition for Vietnamese
por: Le, Thanh-Thien, et al.
Publicado: (2024)
por: Le, Thanh-Thien, et al.
Publicado: (2024)
CantoASR: Prosody-Aware ASR-LALM Collaboration for Low-Resource Cantonese
por: Chen, Dazhong, et al.
Publicado: (2025)
por: Chen, Dazhong, et al.
Publicado: (2025)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
por: Prakash, Jeena, et al.
Publicado: (2025)
por: Prakash, Jeena, et al.
Publicado: (2025)
Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically
por: Shim, Ryan Soh-Eun, et al.
Publicado: (2025)
por: Shim, Ryan Soh-Eun, et al.
Publicado: (2025)
Impact of automatic speech recognition quality on Alzheimer's disease detection from spontaneous speech: a reproducible benchmark study with lexical modeling and statistical validation
por: Samanta, Himadri S
Publicado: (2026)
por: Samanta, Himadri S
Publicado: (2026)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
por: Özyilmaz, Ömer Tarik, et al.
Publicado: (2025)
por: Özyilmaz, Ömer Tarik, et al.
Publicado: (2025)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
por: Vesterbacka, Leonora, et al.
Publicado: (2025)
por: Vesterbacka, Leonora, et al.
Publicado: (2025)
Classification is a RAG problem: A case study on hate speech detection
por: Willats, Richard, et al.
Publicado: (2025)
por: Willats, Richard, et al.
Publicado: (2025)
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
por: Radhakrishnan, Srijith, et al.
Publicado: (2023)
por: Radhakrishnan, Srijith, et al.
Publicado: (2023)
Hypernetworks for Personalizing ASR to Atypical Speech
por: Müller-Eberstein, Max, et al.
Publicado: (2024)
por: Müller-Eberstein, Max, et al.
Publicado: (2024)
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases
por: Zhou, Giulio, et al.
Publicado: (2024)
por: Zhou, Giulio, et al.
Publicado: (2024)
Internalizing ASR with Implicit Chain of Thought for Efficient Speech-to-Speech Conversational LLM
por: Yuen, Robin Shing-Hei, et al.
Publicado: (2024)
por: Yuen, Robin Shing-Hei, et al.
Publicado: (2024)
Distinct Theta Synchrony across Speech Modes: Perceived, Spoken, Whispered, and Imagined
por: Lee, Jung-Sun, et al.
Publicado: (2025)
por: Lee, Jung-Sun, et al.
Publicado: (2025)
Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder
por: Li, Zhengyang, et al.
Publicado: (2026)
por: Li, Zhengyang, et al.
Publicado: (2026)
Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts
por: Altinok, Duygu
Publicado: (2025)
por: Altinok, Duygu
Publicado: (2025)
WhisperNER: Unified Open Named Entity and Speech Recognition
por: Ayache, Gil, et al.
Publicado: (2024)
por: Ayache, Gil, et al.
Publicado: (2024)
Ejemplares similares
-
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
por: Hui, Macarious, et al.
Publicado: (2024) -
WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
por: Chowdhury, Aurchi, et al.
Publicado: (2026) -
A new approach for fine-tuning sentence transformers for intent classification and out-of-scope detection tasks
por: Zhang, Tianyi, et al.
Publicado: (2024) -
Human Latency Conversational Turns for Spoken Avatar Systems
por: Jacoby, Derek, et al.
Publicado: (2024) -
Fine-tuning Whisper for Pashto ASR: strategies and scale
por: Rahman, Hanif
Publicado: (2026)