Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garg, Abhinav, Kim, Jiyeon, Khyalia, Sushil, Kim, Chanwoo, Gowda, Dhananjaya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
von: Malisz, Zofia, et al.
Veröffentlicht: (2024)
von: Malisz, Zofia, et al.
Veröffentlicht: (2024)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
von: Paik, Gio, et al.
Veröffentlicht: (2025)
von: Paik, Gio, et al.
Veröffentlicht: (2025)
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
Transferable speech-to-text large language model alignment module
von: Wu, Boyong, et al.
Veröffentlicht: (2024)
von: Wu, Boyong, et al.
Veröffentlicht: (2024)
Semantic enrichment towards efficient speech representations
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
von: Lyth, Dan, et al.
Veröffentlicht: (2024)
von: Lyth, Dan, et al.
Veröffentlicht: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
FxSearcher: gradient-free text-driven audio transformation
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
Korean aegyo speech shows systematic F1 increase to signal childlike qualities
von: Kim, Ji-eun, et al.
Veröffentlicht: (2026)
von: Kim, Ji-eun, et al.
Veröffentlicht: (2026)
TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data
von: Kim, Seung-Bin, et al.
Veröffentlicht: (2024)
von: Kim, Seung-Bin, et al.
Veröffentlicht: (2024)
An experiment on an automated literature survey of data-driven speech enhancement methods
von: Santos, Arthur dos, et al.
Veröffentlicht: (2023)
von: Santos, Arthur dos, et al.
Veröffentlicht: (2023)
AfriHuBERT: A self-supervised speech representation model for African languages
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
von: Sivasankaran, Sunit, et al.
Veröffentlicht: (2024)
von: Sivasankaran, Sunit, et al.
Veröffentlicht: (2024)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
von: Liu, Oli Danyi, et al.
Veröffentlicht: (2024)
von: Liu, Oli Danyi, et al.
Veröffentlicht: (2024)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
Translating speech with just images
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
Adjust-free adversarial example generation in speech recognition using evolutionary multi-objective optimization under black-box condition
von: Ishida, Shoma, et al.
Veröffentlicht: (2020)
von: Ishida, Shoma, et al.
Veröffentlicht: (2020)
A unified front-end framework for English text-to-speech synthesis
von: Ying, Zelin, et al.
Veröffentlicht: (2023)
von: Ying, Zelin, et al.
Veröffentlicht: (2023)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
An Effective Energy Mask-based Adversarial Evasion Attacks against Misclassification in Speaker Recognition Systems
von: Park, Chanwoo, et al.
Veröffentlicht: (2026)
von: Park, Chanwoo, et al.
Veröffentlicht: (2026)
Self-supervised learning of speech representations with Dutch archival data
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
Can Whisper perform speech-based in-context learning?
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
Direct Punjabi to English speech translation using discrete units
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2025)
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2025)
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
von: Tang, Yun, et al.
Veröffentlicht: (2025)
von: Tang, Yun, et al.
Veröffentlicht: (2025)
Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection
von: Cui, Ziyun, et al.
Veröffentlicht: (2023)
von: Cui, Ziyun, et al.
Veröffentlicht: (2023)
Beyond the binary: Limitations and possibilities of gender-related speech technology research
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2024)
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2023)
Hanprome: Modified Hangeul for Expression of foreign language pronunciation
von: Kim, Wonchan, et al.
Veröffentlicht: (2024)
von: Kim, Wonchan, et al.
Veröffentlicht: (2024)
Exploring the limits of decoder-only models trained on public speech recognition corpora
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025) -
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026) -
Unsupervised lexicon learning from speech is limited by representations rather than clustering
von: Slabbert, Danel, et al.
Veröffentlicht: (2025) -
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024) -
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
von: Malisz, Zofia, et al.
Veröffentlicht: (2024)