Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
Fuente:
arXiv
Salvato in:
| Autori principali: | Nowrin, Sadia, Vertanen, Keith |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Directional Source Separation for Robust Speech Recognition on Smart Glasses
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2025)
di: Chen, Youjun, et al.
Pubblicazione: (2025)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
di: Dutta, Satwik, et al.
Pubblicazione: (2025)
di: Dutta, Satwik, et al.
Pubblicazione: (2025)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
di: Zhao, Zhixian, et al.
Pubblicazione: (2024)
di: Zhao, Zhixian, et al.
Pubblicazione: (2024)
Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces
di: Kuhn, Korbinian, et al.
Pubblicazione: (2025)
di: Kuhn, Korbinian, et al.
Pubblicazione: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
di: Mishra, Ruchik, et al.
Pubblicazione: (2024)
di: Mishra, Ruchik, et al.
Pubblicazione: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
di: Li, Yue, et al.
Pubblicazione: (2024)
di: Li, Yue, et al.
Pubblicazione: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
di: Park, Seohyun, et al.
Pubblicazione: (2025)
di: Park, Seohyun, et al.
Pubblicazione: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
di: Mertes, Silvan, et al.
Pubblicazione: (2024)
di: Mertes, Silvan, et al.
Pubblicazione: (2024)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
di: Wang, Hongbin, et al.
Pubblicazione: (2025)
di: Wang, Hongbin, et al.
Pubblicazione: (2025)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
di: Xiao, Yi, et al.
Pubblicazione: (2022)
di: Xiao, Yi, et al.
Pubblicazione: (2022)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
di: Manukalpa, J. M. Chan Sri, et al.
Pubblicazione: (2025)
di: Manukalpa, J. M. Chan Sri, et al.
Pubblicazione: (2025)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
di: Yu, Luca Jiang-Tao, et al.
Pubblicazione: (2024)
di: Yu, Luca Jiang-Tao, et al.
Pubblicazione: (2024)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
di: Liu, Ailin, et al.
Pubblicazione: (2024)
di: Liu, Ailin, et al.
Pubblicazione: (2024)
Human Feedback Driven Dynamic Speech Emotion Recognition
di: Fedorov, Ilya, et al.
Pubblicazione: (2025)
di: Fedorov, Ilya, et al.
Pubblicazione: (2025)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
di: Benster, Tyler, et al.
Pubblicazione: (2024)
di: Benster, Tyler, et al.
Pubblicazione: (2024)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
di: Chang, Yi, et al.
Pubblicazione: (2024)
di: Chang, Yi, et al.
Pubblicazione: (2024)
MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection
di: Li, Lu, et al.
Pubblicazione: (2025)
di: Li, Lu, et al.
Pubblicazione: (2025)
A Mapping Strategy for Interacting with Latent Audio Synthesis Using Artistic Materials
di: Zheng, Shuoyang, et al.
Pubblicazione: (2024)
di: Zheng, Shuoyang, et al.
Pubblicazione: (2024)
ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
UltrasonicSpheres: Localized, Multi-Channel Sound Spheres Using Off-the-Shelf Speakers and Earables
di: Küttner, Michael, et al.
Pubblicazione: (2025)
di: Küttner, Michael, et al.
Pubblicazione: (2025)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
di: Sankey-Olsen, Cuno, et al.
Pubblicazione: (2025)
di: Sankey-Olsen, Cuno, et al.
Pubblicazione: (2025)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
di: Cai, Zhuojiang, et al.
Pubblicazione: (2024)
di: Cai, Zhuojiang, et al.
Pubblicazione: (2024)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
di: Sharma, Roshan, et al.
Pubblicazione: (2024)
di: Sharma, Roshan, et al.
Pubblicazione: (2024)
Subject Disentanglement Neural Network for Speech Envelope Reconstruction from EEG
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction
di: Hou, Yuanbo, et al.
Pubblicazione: (2024)
di: Hou, Yuanbo, et al.
Pubblicazione: (2024)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
di: Yuan, Kuang, et al.
Pubblicazione: (2025)
di: Yuan, Kuang, et al.
Pubblicazione: (2025)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
di: Kawamura, Kazuki, et al.
Pubblicazione: (2024)
di: Kawamura, Kazuki, et al.
Pubblicazione: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
Enhancing DMI Interactions by Integrating Haptic Feedback for Intricate Vibrato Technique
di: Piao, Ziyue, et al.
Pubblicazione: (2024)
di: Piao, Ziyue, et al.
Pubblicazione: (2024)
A cross-talk robust multichannel VAD model for multiparty agent interactions trained using synthetic re-recordings
di: Han, Hyewon, et al.
Pubblicazione: (2024)
di: Han, Hyewon, et al.
Pubblicazione: (2024)
Interactive Sonification for Health and Energy using ChucK and Unity
di: Zhao, Yichun, et al.
Pubblicazione: (2024)
di: Zhao, Yichun, et al.
Pubblicazione: (2024)
Interfacing with history: Curating with audio augmented objects
di: Cliffe, Laurence
Pubblicazione: (2024)
di: Cliffe, Laurence
Pubblicazione: (2024)
Transhuman Ansambl - Voice Beyond Language
di: Ivsic, Lucija, et al.
Pubblicazione: (2024)
di: Ivsic, Lucija, et al.
Pubblicazione: (2024)
Cervical Auscultation Machine Learning for Dysphagia Assessment
di: Chia, An An, et al.
Pubblicazione: (2024)
di: Chia, An An, et al.
Pubblicazione: (2024)
Springboard, Roadblock or "Crutch"?: How Transgender Users Leverage Voice Changers for Gender Presentation in Social Virtual Reality
di: Povinelli, Kassie, et al.
Pubblicazione: (2024)
di: Povinelli, Kassie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Directional Source Separation for Robust Speech Recognition on Smart Glasses
di: Feng, Tiantian, et al.
Pubblicazione: (2023) -
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2025) -
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
di: Dutta, Satwik, et al.
Pubblicazione: (2025) -
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
di: Zhao, Zhixian, et al.
Pubblicazione: (2024) -
Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces
di: Kuhn, Korbinian, et al.
Pubblicazione: (2025)