A study on weakly-supervised training approaches for phoneme-level pronunciation scoring
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vidal, Jazmín, Ferrer, Luciana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Discovering phoneme-specific critical articulators through a data-driven approach
von: Bandekar, Jesuraj, et al.
Veröffentlicht: (2025)
von: Bandekar, Jesuraj, et al.
Veröffentlicht: (2025)
How phonemes contribute to deep speaker models?
von: Li, Pengqi, et al.
Veröffentlicht: (2024)
von: Li, Pengqi, et al.
Veröffentlicht: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
BabAR: from phoneme recognition to developmental measures of young children's speech production
von: Lavechin, Marvin, et al.
Veröffentlicht: (2026)
von: Lavechin, Marvin, et al.
Veröffentlicht: (2026)
A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception
von: Osses, Alejandro, et al.
Veröffentlicht: (2024)
von: Osses, Alejandro, et al.
Veröffentlicht: (2024)
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
von: Shirahata, Yuma, et al.
Veröffentlicht: (2024)
von: Shirahata, Yuma, et al.
Veröffentlicht: (2024)
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
Study on the Fairness of Speaker Verification Systems on Underrepresented Accents in English
von: Estevez, Mariel, et al.
Veröffentlicht: (2022)
von: Estevez, Mariel, et al.
Veröffentlicht: (2022)
Selecting N-lowest scores for training MOS prediction models
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Hanprome: Modified Hangeul for Expression of foreign language pronunciation
von: Kim, Wonchan, et al.
Veröffentlicht: (2024)
von: Kim, Wonchan, et al.
Veröffentlicht: (2024)
AlignNet: Learning dataset score alignment functions to enable better training of speech quality estimators
von: Pieper, Jaden, et al.
Veröffentlicht: (2024)
von: Pieper, Jaden, et al.
Veröffentlicht: (2024)
Beyond Global Metrics: A Fairness Analysis for Interpretable Voice Disorder Detection Systems
von: Estevez, Mariel, et al.
Veröffentlicht: (2025)
von: Estevez, Mariel, et al.
Veröffentlicht: (2025)
LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
von: Cuervo, Santiago, et al.
Veröffentlicht: (2021)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2021)
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
von: Malisz, Zofia, et al.
Veröffentlicht: (2024)
von: Malisz, Zofia, et al.
Veröffentlicht: (2024)
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
von: Kumar, Anurag, et al.
Veröffentlicht: (2024)
von: Kumar, Anurag, et al.
Veröffentlicht: (2024)
Benchmarking Time-localized Explanations for Audio Classification Models
von: Bolaños, Cecilia, et al.
Veröffentlicht: (2025)
von: Bolaños, Cecilia, et al.
Veröffentlicht: (2025)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
von: Garg, Abhinav, et al.
Veröffentlicht: (2024)
von: Garg, Abhinav, et al.
Veröffentlicht: (2024)
Curriculum learning for self-supervised speaker verification
von: Heo, Hee-Soo, et al.
Veröffentlicht: (2022)
von: Heo, Hee-Soo, et al.
Veröffentlicht: (2022)
Investigating self-supervised features for expressive, multilingual voice conversion
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2025)
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2025)
PiCoGen2: Piano cover generation with transfer learning approach and weakly aligned data
von: Tan, Chih-Pin, et al.
Veröffentlicht: (2024)
von: Tan, Chih-Pin, et al.
Veröffentlicht: (2024)
Align-Consistency: Improving Non-autoregressive and Semi-supervised ASR with Consistency Regularization
von: Huang, Wanting, et al.
Veröffentlicht: (2026)
von: Huang, Wanting, et al.
Veröffentlicht: (2026)
Positive and negative sampling strategies for self-supervised learning on audio-video data
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
von: Pîrlogeanu, Gabriel, et al.
Veröffentlicht: (2026)
von: Pîrlogeanu, Gabriel, et al.
Veröffentlicht: (2026)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
Post-training for Deepfake Speech Detection
von: Ge, Wanying, et al.
Veröffentlicht: (2025)
von: Ge, Wanying, et al.
Veröffentlicht: (2025)
Utilizing synthetic training data for the supervised classification of rat ultrasonic vocalizations
von: Scott, K. Jack, et al.
Veröffentlicht: (2023)
von: Scott, K. Jack, et al.
Veröffentlicht: (2023)
STONE: Self-supervised Tonality Estimator
von: Kong, Yuexuan, et al.
Veröffentlicht: (2024)
von: Kong, Yuexuan, et al.
Veröffentlicht: (2024)
A Pre-training Framework that Encodes Noise Information for Speech Quality Assessment
von: Sultana, Subrina, et al.
Veröffentlicht: (2024)
von: Sultana, Subrina, et al.
Veröffentlicht: (2024)
CardioPHON: Quality assessment and self-supervised pretraining for screening of cardiac function based on phonocardiogram recordings
von: Despotovic, Vladimir, et al.
Veröffentlicht: (2025)
von: Despotovic, Vladimir, et al.
Veröffentlicht: (2025)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model
von: Lavechin, Marvin, et al.
Veröffentlicht: (2025)
von: Lavechin, Marvin, et al.
Veröffentlicht: (2025)
Pre-training Music Classification Models via Music Source Separation
von: Garoufis, Christos, et al.
Veröffentlicht: (2023)
von: Garoufis, Christos, et al.
Veröffentlicht: (2023)
Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training
von: Li, Yiming, et al.
Veröffentlicht: (2024)
von: Li, Yiming, et al.
Veröffentlicht: (2024)
End-to-End Speech Recognition with Pre-trained Masked Language Model
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Discovering phoneme-specific critical articulators through a data-driven approach
von: Bandekar, Jesuraj, et al.
Veröffentlicht: (2025) -
How phonemes contribute to deep speaker models?
von: Li, Pengqi, et al.
Veröffentlicht: (2024) -
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025) -
BabAR: from phoneme recognition to developmental measures of young children's speech production
von: Lavechin, Marvin, et al.
Veröffentlicht: (2026) -
A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception
von: Osses, Alejandro, et al.
Veröffentlicht: (2024)