Discovering phoneme-specific critical articulators through a data-driven approach
Fuente:
arXiv
Salvato in:
| Autori principali: | Bandekar, Jesuraj, Udupa, Sathvik, Ghosh, Prasanta Kumar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
di: Udupa, Sathvik, et al.
Pubblicazione: (2025)
di: Udupa, Sathvik, et al.
Pubblicazione: (2025)
A study on weakly-supervised training approaches for phoneme-level pronunciation scoring
di: Vidal, Jazmín, et al.
Pubblicazione: (2026)
di: Vidal, Jazmín, et al.
Pubblicazione: (2026)
VAANI: Capturing the language landscape for an inclusive digital India
di: Pulikodan, Sujith, et al.
Pubblicazione: (2026)
di: Pulikodan, Sujith, et al.
Pubblicazione: (2026)
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
di: Pulikodan, Sujith, et al.
Pubblicazione: (2025)
di: Pulikodan, Sujith, et al.
Pubblicazione: (2025)
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
di: Shirahata, Yuma, et al.
Pubblicazione: (2024)
di: Shirahata, Yuma, et al.
Pubblicazione: (2024)
How phonemes contribute to deep speaker models?
di: Li, Pengqi, et al.
Pubblicazione: (2024)
di: Li, Pengqi, et al.
Pubblicazione: (2024)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
di: Ma, Te, et al.
Pubblicazione: (2025)
di: Ma, Te, et al.
Pubblicazione: (2025)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
di: Garg, Abhinav, et al.
Pubblicazione: (2024)
di: Garg, Abhinav, et al.
Pubblicazione: (2024)
BabAR: from phoneme recognition to developmental measures of young children's speech production
di: Lavechin, Marvin, et al.
Pubblicazione: (2026)
di: Lavechin, Marvin, et al.
Pubblicazione: (2026)
A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception
di: Osses, Alejandro, et al.
Pubblicazione: (2024)
di: Osses, Alejandro, et al.
Pubblicazione: (2024)
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
di: Zhou, Wangjin, et al.
Pubblicazione: (2024)
di: Zhou, Wangjin, et al.
Pubblicazione: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
di: Dong, Lukuang, et al.
Pubblicazione: (2026)
di: Dong, Lukuang, et al.
Pubblicazione: (2026)
Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
di: Cuervo, Santiago, et al.
Pubblicazione: (2021)
di: Cuervo, Santiago, et al.
Pubblicazione: (2021)
Role of the Pretraining and the Adaptation data sizes for low-resource real-time MRI video segmentation
di: Tholan, Masoud Thajudeen, et al.
Pubblicazione: (2025)
di: Tholan, Masoud Thajudeen, et al.
Pubblicazione: (2025)
Bottleneck Transformer-Based Approach for Improved Automatic STOI Score Prediction
di: Amartyaveer, et al.
Pubblicazione: (2026)
di: Amartyaveer, et al.
Pubblicazione: (2026)
Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?
di: Dutta, Bikash, et al.
Pubblicazione: (2025)
di: Dutta, Bikash, et al.
Pubblicazione: (2025)
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
di: Malisz, Zofia, et al.
Pubblicazione: (2024)
di: Malisz, Zofia, et al.
Pubblicazione: (2024)
Unmasking real-world audio deepfakes: A data-centric approach
di: Combei, David, et al.
Pubblicazione: (2025)
di: Combei, David, et al.
Pubblicazione: (2025)
Learning to Discover: A Generalized Framework for Raga Identification without Forgetting
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
Improving acoustic drone detection generalization through pretraining and data augmentation
di: Reuter, Paul M., et al.
Pubblicazione: (2026)
di: Reuter, Paul M., et al.
Pubblicazione: (2026)
Complete reconstruction of the tongue contour through acoustic to articulatory inversion using real-time MRI data
di: Azzouz, Sofiane, et al.
Pubblicazione: (2024)
di: Azzouz, Sofiane, et al.
Pubblicazione: (2024)
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
di: Kumar, Sonal, et al.
Pubblicazione: (2025)
di: Kumar, Sonal, et al.
Pubblicazione: (2025)
ProSE: Diffusion Priors for Speech Enhancement
di: Kumar, Sonal, et al.
Pubblicazione: (2025)
di: Kumar, Sonal, et al.
Pubblicazione: (2025)
Deep, data-driven modeling of room acoustics: literature review and research perspectives
di: van Waterschoot, Toon
Pubblicazione: (2025)
di: van Waterschoot, Toon
Pubblicazione: (2025)
Exploring the anatomy of articulation rate in spontaneous English speech: relationships between utterance length effects and social factors
di: Tanner, James, et al.
Pubblicazione: (2024)
di: Tanner, James, et al.
Pubblicazione: (2024)
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
di: Kumar, Anurag, et al.
Pubblicazione: (2024)
di: Kumar, Anurag, et al.
Pubblicazione: (2024)
High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
di: Banerjee, Sourav, et al.
Pubblicazione: (2024)
di: Banerjee, Sourav, et al.
Pubblicazione: (2024)
Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
di: Cumlin, Fredrik, et al.
Pubblicazione: (2025)
di: Cumlin, Fredrik, et al.
Pubblicazione: (2025)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
di: Seth, Ashish, et al.
Pubblicazione: (2024)
di: Seth, Ashish, et al.
Pubblicazione: (2024)
Improving Stereo 3D Sound Event Localization and Detection: Perceptual Features, Stereo-specific Data Augmentation, and Distance Normalization
di: Yeow, Jun-Wei, et al.
Pubblicazione: (2025)
di: Yeow, Jun-Wei, et al.
Pubblicazione: (2025)
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024)
di: Li, Mohan, et al.
Pubblicazione: (2024)
QiandaoEar22: A high quality noise dataset for identifying specific ship from multiple underwater acoustic targets using ship-radiated noise
di: Du, Xiaoyang, et al.
Pubblicazione: (2024)
di: Du, Xiaoyang, et al.
Pubblicazione: (2024)
Prompt-driven Target Speech Diarization
di: Jiang, Yidi, et al.
Pubblicazione: (2023)
di: Jiang, Yidi, et al.
Pubblicazione: (2023)
Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction
di: Du, Hui-Peng, et al.
Pubblicazione: (2026)
di: Du, Hui-Peng, et al.
Pubblicazione: (2026)
Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
di: Yusuf, Bolaji, et al.
Pubblicazione: (2024)
di: Yusuf, Bolaji, et al.
Pubblicazione: (2024)
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
di: C, Shiva Kumar, et al.
Pubblicazione: (2025)
di: C, Shiva Kumar, et al.
Pubblicazione: (2025)
PiCoGen2: Piano cover generation with transfer learning approach and weakly aligned data
di: Tan, Chih-Pin, et al.
Pubblicazione: (2024)
di: Tan, Chih-Pin, et al.
Pubblicazione: (2024)
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
di: Zhao, Xiutian, et al.
Pubblicazione: (2026)
di: Zhao, Xiutian, et al.
Pubblicazione: (2026)
MUSHRA-1S: A scalable and sensitive test approach for evaluating top-tier speech processing systems
di: Lechler, Laura, et al.
Pubblicazione: (2025)
di: Lechler, Laura, et al.
Pubblicazione: (2025)
Towards Improved Speech Recognition through Optimized Synthetic Data Generation
di: Perrin, Yanis, et al.
Pubblicazione: (2025)
di: Perrin, Yanis, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
di: Udupa, Sathvik, et al.
Pubblicazione: (2025) -
A study on weakly-supervised training approaches for phoneme-level pronunciation scoring
di: Vidal, Jazmín, et al.
Pubblicazione: (2026) -
VAANI: Capturing the language landscape for an inclusive digital India
di: Pulikodan, Sujith, et al.
Pubblicazione: (2026) -
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
di: Pulikodan, Sujith, et al.
Pubblicazione: (2025) -
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
di: Shirahata, Yuma, et al.
Pubblicazione: (2024)