Disentangling segmental and prosodic factors to non-native speech comprehensibility
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Quamer, Waris, Gutierrez-Osuna, Ricardo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
End-to-end streaming model for low-latency speech anonymization
von: Quamer, Waris, et al.
Veröffentlicht: (2024)
von: Quamer, Waris, et al.
Veröffentlicht: (2024)
DarkStream: real-time speech anonymization with low latency
von: Quamer, Waris, et al.
Veröffentlicht: (2025)
von: Quamer, Waris, et al.
Veröffentlicht: (2025)
PHONOS: PHOnetic Neutralization for Online Streaming Applications
von: Quamer, Waris, et al.
Veröffentlicht: (2026)
von: Quamer, Waris, et al.
Veröffentlicht: (2026)
TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
von: Quamer, Waris, et al.
Veröffentlicht: (2026)
von: Quamer, Waris, et al.
Veröffentlicht: (2026)
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
von: Shirahata, Yuma, et al.
Veröffentlicht: (2024)
von: Shirahata, Yuma, et al.
Veröffentlicht: (2024)
Prominence-aware automatic speech recognition for conversational speech
von: Linke, Julian, et al.
Veröffentlicht: (2025)
von: Linke, Julian, et al.
Veröffentlicht: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
On the reliability of feature attribution methods for speech classification
von: Shen, Gaofei, et al.
Veröffentlicht: (2025)
von: Shen, Gaofei, et al.
Veröffentlicht: (2025)
Forensic deepfake audio detection using segmental speech features
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
The mutual exclusivity bias of bilingual visually grounded speech models
von: Oneata, Dan, et al.
Veröffentlicht: (2025)
von: Oneata, Dan, et al.
Veröffentlicht: (2025)
Non-native Children's Automatic Speech Assessment Challenge (NOCASA)
von: Getman, Yaroslav, et al.
Veröffentlicht: (2025)
von: Getman, Yaroslav, et al.
Veröffentlicht: (2025)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
Translating speech with just images
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
Exploring the anatomy of articulation rate in spontaneous English speech: relationships between utterance length effects and social factors
von: Tanner, James, et al.
Veröffentlicht: (2024)
von: Tanner, James, et al.
Veröffentlicht: (2024)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
von: Ogun, Sewade, et al.
Veröffentlicht: (2024)
von: Ogun, Sewade, et al.
Veröffentlicht: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
von: San, Nay, et al.
Veröffentlicht: (2024)
von: San, Nay, et al.
Veröffentlicht: (2024)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
von: Gubian, Michele, et al.
Veröffentlicht: (2025)
von: Gubian, Michele, et al.
Veröffentlicht: (2025)
Semantic enrichment towards efficient speech representations
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2025)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2025)
Can Whisper perform speech-based in-context learning?
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
Direct Punjabi to English speech translation using discrete units
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
Transferable speech-to-text large language model alignment module
von: Wu, Boyong, et al.
Veröffentlicht: (2024)
von: Wu, Boyong, et al.
Veröffentlicht: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
von: Lyth, Dan, et al.
Veröffentlicht: (2024)
von: Lyth, Dan, et al.
Veröffentlicht: (2024)
Beyond the binary: Limitations and possibilities of gender-related speech technology research
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2024)
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2024)
Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection
von: Cui, Ziyun, et al.
Veröffentlicht: (2023)
von: Cui, Ziyun, et al.
Veröffentlicht: (2023)
An experiment on an automated literature survey of data-driven speech enhancement methods
von: Santos, Arthur dos, et al.
Veröffentlicht: (2023)
von: Santos, Arthur dos, et al.
Veröffentlicht: (2023)
Convoifilter: A case study of doing cocktail party speech recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
Exploring the limits of decoder-only models trained on public speech recognition corpora
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
von: Garg, Abhinav, et al.
Veröffentlicht: (2024)
von: Garg, Abhinav, et al.
Veröffentlicht: (2024)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
von: Sedukhin, Oleg, et al.
Veröffentlicht: (2026)
von: Sedukhin, Oleg, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
End-to-end streaming model for low-latency speech anonymization
von: Quamer, Waris, et al.
Veröffentlicht: (2024) -
DarkStream: real-time speech anonymization with low latency
von: Quamer, Waris, et al.
Veröffentlicht: (2025) -
PHONOS: PHOnetic Neutralization for Online Streaming Applications
von: Quamer, Waris, et al.
Veröffentlicht: (2026) -
TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
von: Quamer, Waris, et al.
Veröffentlicht: (2026) -
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024)