Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Tu, Huu Tuong, Khanh, Ha Viet, Dat, Tran Tien, Huan, Vu, Van Luong, Thien, Cuong, Nguyen Tien, Trang, Nguyen Thi Thu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
by: Tu, Huu Tuong, et al.
Published: (2025)
by: Tu, Huu Tuong, et al.
Published: (2025)
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
by: Vu, Hoang Long, et al.
Published: (2024)
by: Vu, Hoang Long, et al.
Published: (2024)
Zero-Shot Text-to-Speech for Vietnamese
by: Vu, Thi, et al.
Published: (2025)
by: Vu, Thi, et al.
Published: (2025)
PhoWhisper: Automatic Speech Recognition for Vietnamese
by: Le, Thanh-Thien, et al.
Published: (2024)
by: Le, Thanh-Thien, et al.
Published: (2024)
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
by: Hoang, Long-Vu, et al.
Published: (2025)
by: Hoang, Long-Vu, et al.
Published: (2025)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
by: Nguyen, Tuan, et al.
Published: (2025)
by: Nguyen, Tuan, et al.
Published: (2025)
Wanna hear your voice? A sample is all we need!
by: Pham, The Hieu, et al.
Published: (2024)
by: Pham, The Hieu, et al.
Published: (2024)
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders
by: Lam, Phat, et al.
Published: (2024)
by: Lam, Phat, et al.
Published: (2024)
Pitch-Aware RNN-T for Mandarin Chinese Mispronunciation Detection and Diagnosis
by: Wang, Xintong, et al.
Published: (2024)
by: Wang, Xintong, et al.
Published: (2024)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
by: Nguyen, Tuan, et al.
Published: (2025)
by: Nguyen, Tuan, et al.
Published: (2025)
Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis
by: Geng, Haopeng, et al.
Published: (2026)
by: Geng, Haopeng, et al.
Published: (2026)
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
by: Ho, Luong, et al.
Published: (2025)
by: Ho, Luong, et al.
Published: (2025)
Detecting abnormal heart sound using mobile phones and on-device IConNet
by: Vu, Linh, et al.
Published: (2024)
by: Vu, Linh, et al.
Published: (2024)
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
by: Nguyen, Tuan, et al.
Published: (2025)
by: Nguyen, Tuan, et al.
Published: (2025)
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
by: Phuong, Tuan Dat, et al.
Published: (2025)
by: Phuong, Tuan Dat, et al.
Published: (2025)
Continuous Learning of Transformer-based Audio Deepfake Detection
by: Le, Tuan Duy Nguyen, et al.
Published: (2024)
by: Le, Tuan Duy Nguyen, et al.
Published: (2024)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
by: Pham, The Hieu, et al.
Published: (2025)
by: Pham, The Hieu, et al.
Published: (2025)
Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge
by: Parikh, Aditya Kamlesh, et al.
Published: (2025)
by: Parikh, Aditya Kamlesh, et al.
Published: (2025)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
by: Pham, Lam, et al.
Published: (2024)
by: Pham, Lam, et al.
Published: (2024)
Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
by: Phan, Nhan, et al.
Published: (2025)
by: Phan, Nhan, et al.
Published: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
by: Le, Khanh, et al.
Published: (2025)
by: Le, Khanh, et al.
Published: (2025)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
by: Le, Khanh, et al.
Published: (2025)
by: Le, Khanh, et al.
Published: (2025)
Evaluating Logit-Based GOP Scores for Mispronunciation Detection
by: Parikh, Aditya Kamlesh, et al.
Published: (2025)
by: Parikh, Aditya Kamlesh, et al.
Published: (2025)
An approach to hummed-tune and song sequences matching
by: Pham, Loc Bao, et al.
Published: (2024)
by: Pham, Loc Bao, et al.
Published: (2024)
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
TokenSE: a Mamba-based discrete token speech enhancement framework for cochlear implants
by: Chiang, Hsin-Tien, et al.
Published: (2026)
by: Chiang, Hsin-Tien, et al.
Published: (2026)
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
by: Nguyen, Tan Dat, et al.
Published: (2026)
by: Nguyen, Tan Dat, et al.
Published: (2026)
XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
by: Dat, Phuong Tuan, et al.
Published: (2025)
by: Dat, Phuong Tuan, et al.
Published: (2025)
Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources
by: Nguyen, Binh Thien, et al.
Published: (2026)
by: Nguyen, Binh Thien, et al.
Published: (2026)
Revisiting Deep Audio-Text Retrieval Through the Lens of Transportation
by: Luong, Manh, et al.
Published: (2024)
by: Luong, Manh, et al.
Published: (2024)
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
by: Pham, Long-Khanh, et al.
Published: (2025)
by: Pham, Long-Khanh, et al.
Published: (2025)
Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
by: Lee, Kyowoon, et al.
Published: (2025)
by: Lee, Kyowoon, et al.
Published: (2025)
Toward end-to-end interpretable convolutional neural networks for waveform signals
by: Vu, Linh, et al.
Published: (2024)
by: Vu, Linh, et al.
Published: (2024)
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
by: Roll, Nathan, et al.
Published: (2025)
by: Roll, Nathan, et al.
Published: (2025)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
by: Tsai, Kai-Chen, et al.
Published: (2026)
by: Tsai, Kai-Chen, et al.
Published: (2026)
An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding
by: Su, Tianhui, et al.
Published: (2026)
by: Su, Tianhui, et al.
Published: (2026)
The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines using Deep Learning Based Model
by: Nguyen, Tin, et al.
Published: (2024)
by: Nguyen, Tin, et al.
Published: (2024)
EffiFusion-GAN: Efficient Fusion Generative Adversarial Network for Speech Enhancement
by: Wen, Bin, et al.
Published: (2025)
by: Wen, Bin, et al.
Published: (2025)
RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations
by: Luong, Hieu-Thi, et al.
Published: (2026)
by: Luong, Hieu-Thi, et al.
Published: (2026)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
by: Le-Duc, Khai, et al.
Published: (2025)
by: Le-Duc, Khai, et al.
Published: (2025)
Similar Items
-
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
by: Tu, Huu Tuong, et al.
Published: (2025) -
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
by: Vu, Hoang Long, et al.
Published: (2024) -
Zero-Shot Text-to-Speech for Vietnamese
by: Vu, Thi, et al.
Published: (2025) -
PhoWhisper: Automatic Speech Recognition for Vietnamese
by: Le, Thanh-Thien, et al.
Published: (2024) -
Acoustic scattering AI for non-invasive object classifications: A case study on hair assessment
by: Hoang, Long-Vu, et al.
Published: (2025)