Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Shuiyun, Kong, Yuxiang, Guo, Pengcheng, Zhuang, Weiji, Gao, Peng, Wang, Yujun, Xie, Lei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge
por: Wang, Shiyao, et al.
Publicado: (2024)
por: Wang, Shiyao, et al.
Publicado: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
por: Guo, Zhao, et al.
Publicado: (2025)
por: Guo, Zhao, et al.
Publicado: (2025)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
CDSD: Chinese Dysarthria Speech Database
por: Wang, Yan, et al.
Publicado: (2023)
por: Wang, Yan, et al.
Publicado: (2023)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
por: Wang, Haoxu, et al.
Publicado: (2024)
por: Wang, Haoxu, et al.
Publicado: (2024)
End-to-End User-Defined Keyword Spotting using Shifted Delta Coefficients
por: V, Kesavaraj, et al.
Publicado: (2024)
por: V, Kesavaraj, et al.
Publicado: (2024)
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
por: Peng, Jing, et al.
Publicado: (2026)
por: Peng, Jing, et al.
Publicado: (2026)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
por: Lin, Wan, et al.
Publicado: (2024)
por: Lin, Wan, et al.
Publicado: (2024)
An End-to-End Approach for Chord-Conditioned Song Generation
por: Gao, Shuochen, et al.
Publicado: (2024)
por: Gao, Shuochen, et al.
Publicado: (2024)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
por: Guimarães, Heitor R., et al.
Publicado: (2024)
por: Guimarães, Heitor R., et al.
Publicado: (2024)
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
por: Zhuang, Zhuoran, et al.
Publicado: (2026)
por: Zhuang, Zhuoran, et al.
Publicado: (2026)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
por: Ghane, Mohsen, et al.
Publicado: (2025)
por: Ghane, Mohsen, et al.
Publicado: (2025)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
por: Wang, Zhichao, et al.
Publicado: (2024)
por: Wang, Zhichao, et al.
Publicado: (2024)
End-to-End Diarization utilizing Attractor Deep Clustering
por: Palzer, David, et al.
Publicado: (2025)
por: Palzer, David, et al.
Publicado: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
por: Zhao, Qiuming, et al.
Publicado: (2024)
por: Zhao, Qiuming, et al.
Publicado: (2024)
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
por: You, Zhenghai, et al.
Publicado: (2025)
por: You, Zhenghai, et al.
Publicado: (2025)
Low-latency auditory spatial attention detection based on spectro-spatial features from EEG
por: Cai, Siqi, et al.
Publicado: (2021)
por: Cai, Siqi, et al.
Publicado: (2021)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
por: Plaquet, Alexis, et al.
Publicado: (2025)
por: Plaquet, Alexis, et al.
Publicado: (2025)
Implementation and Applications of WakeWords Integrated with Speaker Recognition: A Case Study
por: Filho, Alexandre Costa Ferro, et al.
Publicado: (2024)
por: Filho, Alexandre Costa Ferro, et al.
Publicado: (2024)
WSCoach: Wearable Real-time Auditory Feedback for Reducing Unwanted Words in Daily Communication
por: Youpeng, Zhang, et al.
Publicado: (2025)
por: Youpeng, Zhang, et al.
Publicado: (2025)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
por: Kang, Wonjune, et al.
Publicado: (2022)
por: Kang, Wonjune, et al.
Publicado: (2022)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
por: Li, Shaojun, et al.
Publicado: (2024)
por: Li, Shaojun, et al.
Publicado: (2024)
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors
por: Landini, Federico, et al.
Publicado: (2023)
por: Landini, Federico, et al.
Publicado: (2023)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
por: Cai, Zhuojiang, et al.
Publicado: (2024)
por: Cai, Zhuojiang, et al.
Publicado: (2024)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
por: Zhou, Junzuo, et al.
Publicado: (2024)
por: Zhou, Junzuo, et al.
Publicado: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
por: Ahmad, Hawraz A., et al.
Publicado: (2024)
por: Ahmad, Hawraz A., et al.
Publicado: (2024)
End-to-End Amp Modeling: From Data to Controllable Guitar Amplifier Models
por: Juvela, Lauri, et al.
Publicado: (2024)
por: Juvela, Lauri, et al.
Publicado: (2024)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
por: Kushwaha, Saksham Singh, et al.
Publicado: (2024)
por: Kushwaha, Saksham Singh, et al.
Publicado: (2024)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
por: Xu, Anfeng, et al.
Publicado: (2026)
por: Xu, Anfeng, et al.
Publicado: (2026)
Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
por: Yu, Chin-Yun, et al.
Publicado: (2024)
por: Yu, Chin-Yun, et al.
Publicado: (2024)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
por: Raissi, Tina, et al.
Publicado: (2025)
por: Raissi, Tina, et al.
Publicado: (2025)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
por: Zhao, Qiuming, et al.
Publicado: (2024)
por: Zhao, Qiuming, et al.
Publicado: (2024)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
por: Zhao, Zhixian, et al.
Publicado: (2024)
por: Zhao, Zhixian, et al.
Publicado: (2024)
RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
por: Chen, Yujie, et al.
Publicado: (2024)
por: Chen, Yujie, et al.
Publicado: (2024)
Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?
por: Zhang, Lin, et al.
Publicado: (2024)
por: Zhang, Lin, et al.
Publicado: (2024)
audio2chart: End to End Audio Transcription into playable Guitar Hero charts
por: Tripodi, Riccardo
Publicado: (2025)
por: Tripodi, Riccardo
Publicado: (2025)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
por: Guo, Yinlin, et al.
Publicado: (2024)
por: Guo, Yinlin, et al.
Publicado: (2024)
Prototype: A Keyword Spotting-Based Intelligent Audio SoC for IoT
por: Liang, Huihong, et al.
Publicado: (2025)
por: Liang, Huihong, et al.
Publicado: (2025)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
por: Du, Zongyang, et al.
Publicado: (2024)
por: Du, Zongyang, et al.
Publicado: (2024)
LS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction
por: Liang, Di, et al.
Publicado: (2024)
por: Liang, Di, et al.
Publicado: (2024)
Ejemplares similares
-
PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge
por: Wang, Shiyao, et al.
Publicado: (2024) -
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
por: Guo, Zhao, et al.
Publicado: (2025) -
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
por: Wu, Haibin, et al.
Publicado: (2024) -
CDSD: Chinese Dysarthria Speech Database
por: Wang, Yan, et al.
Publicado: (2023) -
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
por: Wang, Haoxu, et al.
Publicado: (2024)