Segmentation-free Goodness of Pronunciation
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Xinwei, Fan, Zijian, Svendsen, Torbjørn, Salvi, Giampiero |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection
by: Cao, Xinwei, et al.
Published: (2026)
by: Cao, Xinwei, et al.
Published: (2026)
Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
by: Yusuyin, Saierdaer, et al.
Published: (2025)
by: Yusuyin, Saierdaer, et al.
Published: (2025)
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
Non-native Children's Automatic Speech Assessment Challenge (NOCASA)
by: Getman, Yaroslav, et al.
Published: (2025)
by: Getman, Yaroslav, et al.
Published: (2025)
K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function
by: Li, Shuhe, et al.
Published: (2025)
by: Li, Shuhe, et al.
Published: (2025)
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
by: Fan, Xiaoran, et al.
Published: (2025)
by: Fan, Xiaoran, et al.
Published: (2025)
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
by: La Quatra, Moreno, et al.
Published: (2024)
by: La Quatra, Moreno, et al.
Published: (2024)
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
by: Abdelfattah, Abdullah, et al.
Published: (2025)
by: Abdelfattah, Abdullah, et al.
Published: (2025)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
by: Djanibekov, Amirbek, et al.
Published: (2026)
by: Djanibekov, Amirbek, et al.
Published: (2026)
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
by: Kumar, Shashi, et al.
Published: (2024)
by: Kumar, Shashi, et al.
Published: (2024)
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
by: Fu, Li, et al.
Published: (2025)
by: Fu, Li, et al.
Published: (2025)
Pronunciation Assessment with Multi-modal Large Language Models
by: Fu, Kaiqi, et al.
Published: (2024)
by: Fu, Kaiqi, et al.
Published: (2024)
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
by: Zhao, Rui, et al.
Published: (2024)
by: Zhao, Rui, et al.
Published: (2024)
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
by: Cao, Di, et al.
Published: (2026)
by: Cao, Di, et al.
Published: (2026)
Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing
by: Cheng, Gaofeng, et al.
Published: (2025)
by: Cheng, Gaofeng, et al.
Published: (2025)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
by: Yang, Guanrou, et al.
Published: (2024)
by: Yang, Guanrou, et al.
Published: (2024)
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects
by: Yang, Sicheng, et al.
Published: (2026)
by: Yang, Sicheng, et al.
Published: (2026)
MuFFIN: Multifaceted Pronunciation Feedback Model with Interactive Hierarchical Neural Modeling
by: Yan, Bi-Cheng, et al.
Published: (2025)
by: Yan, Bi-Cheng, et al.
Published: (2025)
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
by: Yu, Wenyi, et al.
Published: (2024)
by: Yu, Wenyi, et al.
Published: (2024)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
by: Polák, Peter, et al.
Published: (2023)
by: Polák, Peter, et al.
Published: (2023)
Pretraining Large Brain Language Model for Active BCI: Silent Speech
by: Zhou, Jinzhao, et al.
Published: (2025)
by: Zhou, Jinzhao, et al.
Published: (2025)
CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
by: Yang, Guanrou, et al.
Published: (2025)
by: Yang, Guanrou, et al.
Published: (2025)
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
by: Liu, Changsong, et al.
Published: (2025)
by: Liu, Changsong, et al.
Published: (2025)
Segmental Attention Decoding With Long Form Acoustic Encodings
by: Swietojanski, Pawel, et al.
Published: (2025)
by: Swietojanski, Pawel, et al.
Published: (2025)
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
by: Valdivia, Andrew, et al.
Published: (2025)
by: Valdivia, Andrew, et al.
Published: (2025)
Towards Unsupervised Speech Recognition Without Pronunciation Models
by: Ni, Junrui, et al.
Published: (2024)
by: Ni, Junrui, et al.
Published: (2024)
Zero-Shot Text-to-Speech as Golden Speech Generator: A Systematic Framework and its Applicability in Automatic Pronunciation Assessment
by: Lo, Tien-Hong, et al.
Published: (2024)
by: Lo, Tien-Hong, et al.
Published: (2024)
Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
by: Zhang, Shucong, et al.
Published: (2025)
by: Zhang, Shucong, et al.
Published: (2025)
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
by: Wallbridge, Sarenne, et al.
Published: (2025)
by: Wallbridge, Sarenne, et al.
Published: (2025)
Audio-Aware Large Language Models as Judges for Speaking Styles
by: Chiang, Cheng-Han, et al.
Published: (2025)
by: Chiang, Cheng-Han, et al.
Published: (2025)
Multimodal Proposal for an AI-Based Tool to Increase Cross-Assessment of Messages
by: Castro, Alejandro Álvarez, et al.
Published: (2025)
by: Castro, Alejandro Álvarez, et al.
Published: (2025)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
by: Fong, Seraphina, et al.
Published: (2025)
by: Fong, Seraphina, et al.
Published: (2025)
Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMs
by: Cuervo, Santiago, et al.
Published: (2025)
by: Cuervo, Santiago, et al.
Published: (2025)
ViDove: A Translation Agent System with Multimodal Context and Memory-Augmented Reasoning
by: Lu, Yichen, et al.
Published: (2025)
by: Lu, Yichen, et al.
Published: (2025)
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
by: Lee, Jihwan, et al.
Published: (2025)
by: Lee, Jihwan, et al.
Published: (2025)
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
by: Labrak, Yanis, et al.
Published: (2025)
by: Labrak, Yanis, et al.
Published: (2025)
Similar Items
-
Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection
by: Cao, Xinwei, et al.
Published: (2026) -
Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
by: Yusuyin, Saierdaer, et al.
Published: (2025) -
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
by: Do, Heejin, et al.
Published: (2024) -
Non-native Children's Automatic Speech Assessment Challenge (NOCASA)
by: Getman, Yaroslav, et al.
Published: (2025) -
K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function
by: Li, Shuhe, et al.
Published: (2025)