Visual-Aware Speech Recognition for Noisy Scenarios
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Balaji, Lakshmipathi, Singla, Karan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
von: Wills, Simone, et al.
Veröffentlicht: (2023)
von: Wills, Simone, et al.
Veröffentlicht: (2023)
Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
A Large-Scale Evaluation of Speech Foundation Models
von: Yang, Shu-wen, et al.
Veröffentlicht: (2024)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2024)
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
WST-X Series: Wavelet Scattering Transform for Interpretable Speech Deepfake Detection
von: Xuan, Xi, et al.
Veröffentlicht: (2026)
von: Xuan, Xi, et al.
Veröffentlicht: (2026)
WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection
von: Xuan, Xi, et al.
Veröffentlicht: (2025)
von: Xuan, Xi, et al.
Veröffentlicht: (2025)
RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
von: Goswami, Mandip
Veröffentlicht: (2026)
von: Goswami, Mandip
Veröffentlicht: (2026)
Hybrid Deep Learning and Signal Processing for Arabic Dialect Recognition in Low-Resource Settings
von: Al-Shwayyat, Ghazal, et al.
Veröffentlicht: (2025)
von: Al-Shwayyat, Ghazal, et al.
Veröffentlicht: (2025)
Semantic Communications for Speech Recognition
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios with Synthetic Visual Data
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
von: Storey, Edward, et al.
Veröffentlicht: (2025)
von: Storey, Edward, et al.
Veröffentlicht: (2025)
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech
von: Mohammad, Mir Sayeed, et al.
Veröffentlicht: (2024)
von: Mohammad, Mir Sayeed, et al.
Veröffentlicht: (2024)
A Computational Approach to Analyzing Disrupted Language in Schizophrenia: Integrating Surprisal and Coherence Measures
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
Neuro2Semantic: A Transfer Learning Framework for Semantic Reconstruction of Continuous Language from Human Intracranial EEG
von: Shams, Siavash, et al.
Veröffentlicht: (2025)
von: Shams, Siavash, et al.
Veröffentlicht: (2025)
MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
von: Khurana, Sameer, et al.
Veröffentlicht: (2023)
von: Khurana, Sameer, et al.
Veröffentlicht: (2023)
Reading Miscue Detection in Primary School through Automatic Speech Recognition
von: Gao, Lingyun, et al.
Veröffentlicht: (2024)
von: Gao, Lingyun, et al.
Veröffentlicht: (2024)
Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios
von: Jiang, Ya, et al.
Veröffentlicht: (2024)
von: Jiang, Ya, et al.
Veröffentlicht: (2024)
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
A Study on Speech Assessment with Visual Cues
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
SpeechMLC: Speech Multi-label Classification
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations
von: Dutta, Soumya, et al.
Veröffentlicht: (2026)
von: Dutta, Soumya, et al.
Veröffentlicht: (2026)
Exploring Dynamic Parameters for Vietnamese Gender-Independent ASR
von: Leang, Sotheara, et al.
Veröffentlicht: (2025)
von: Leang, Sotheara, et al.
Veröffentlicht: (2025)
Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses
von: Gómez-Zaragozá, Lucía, et al.
Veröffentlicht: (2023)
von: Gómez-Zaragozá, Lucía, et al.
Veröffentlicht: (2023)
Automatic Assessment of Oral Reading Accuracy for Reading Diagnostics
von: Molenaar, Bo, et al.
Veröffentlicht: (2023)
von: Molenaar, Bo, et al.
Veröffentlicht: (2023)
Is Attention always needed? A Case Study on Language Identification from Speech
von: Mandal, Atanu, et al.
Veröffentlicht: (2021)
von: Mandal, Atanu, et al.
Veröffentlicht: (2021)
Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection
von: Mandal, Atanu, et al.
Veröffentlicht: (2024)
von: Mandal, Atanu, et al.
Veröffentlicht: (2024)
Benchmarking Audio Deepfake Detection Robustness in Real-world Communication Scenarios
von: Shi, Haohan, et al.
Veröffentlicht: (2025)
von: Shi, Haohan, et al.
Veröffentlicht: (2025)
Target Speaker Selection for Neural Network Beamforming in Multi-Speaker Scenarios
von: Fiorio, Luan Vinícius, et al.
Veröffentlicht: (2025)
von: Fiorio, Luan Vinícius, et al.
Veröffentlicht: (2025)
A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling
von: Marmor, Yanir, et al.
Veröffentlicht: (2026)
von: Marmor, Yanir, et al.
Veröffentlicht: (2026)
Binaural Localization Model for Speech in Noise
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
Speech-Based Prioritization for Schizophrenia Intervention
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
Prompt-driven Target Speech Diarization
von: Jiang, Yidi, et al.
Veröffentlicht: (2023)
von: Jiang, Yidi, et al.
Veröffentlicht: (2023)
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
von: Fu, Li, et al.
Veröffentlicht: (2025)
von: Fu, Li, et al.
Veröffentlicht: (2025)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Streaming Speech-to-Confusion Network Speech Recognition
von: Filimonov, Denis, et al.
Veröffentlicht: (2023)
von: Filimonov, Denis, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
von: Wills, Simone, et al.
Veröffentlicht: (2023) -
Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation
von: Kim, Minsu, et al.
Veröffentlicht: (2023) -
A Large-Scale Evaluation of Speech Foundation Models
von: Yang, Shu-wen, et al.
Veröffentlicht: (2024) -
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024) -
WST-X Series: Wavelet Scattering Transform for Interpretable Speech Deepfake Detection
von: Xuan, Xi, et al.
Veröffentlicht: (2026)