Multimodal Input Aids a Bayesian Model of Phonetic Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhi, Sophia, Levy, Roger P., Meylan, Stephan C. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Phonetic Segmentation of the UCLA Phonetics Lab Archive
di: Chodroff, Eleanor, et al.
Pubblicazione: (2024)
di: Chodroff, Eleanor, et al.
Pubblicazione: (2024)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
di: Zhou, Kun, et al.
Pubblicazione: (2024)
di: Zhou, Kun, et al.
Pubblicazione: (2024)
Phonetic Error Analysis of Raw Waveform Acoustic Models with Parametric and Non-Parametric CNNs
di: Loweimi, Erfan, et al.
Pubblicazione: (2024)
di: Loweimi, Erfan, et al.
Pubblicazione: (2024)
A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs
di: Šegedin, Bruno Ferenc
Pubblicazione: (2025)
di: Šegedin, Bruno Ferenc
Pubblicazione: (2025)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
di: Yusuyin, Saierdaer, et al.
Pubblicazione: (2024)
di: Yusuyin, Saierdaer, et al.
Pubblicazione: (2024)
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
di: Yamashita, Natsuo, et al.
Pubblicazione: (2025)
di: Yamashita, Natsuo, et al.
Pubblicazione: (2025)
(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test
di: Bleeck, Stefan
Pubblicazione: (2025)
di: Bleeck, Stefan
Pubblicazione: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
di: Har-Tuv, Nadav, et al.
Pubblicazione: (2025)
di: Har-Tuv, Nadav, et al.
Pubblicazione: (2025)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
di: Li, Xingyuan, et al.
Pubblicazione: (2024)
di: Li, Xingyuan, et al.
Pubblicazione: (2024)
ISPA: Inter-Species Phonetic Alphabet for Transcribing Animal Sounds
di: Hagiwara, Masato, et al.
Pubblicazione: (2024)
di: Hagiwara, Masato, et al.
Pubblicazione: (2024)
Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
di: Choi, Kwanghee, et al.
Pubblicazione: (2026)
di: Choi, Kwanghee, et al.
Pubblicazione: (2026)
PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction
di: Li, Shufan, et al.
Pubblicazione: (2025)
di: Li, Shufan, et al.
Pubblicazione: (2025)
The Mason-Alberta Phonetic Segmenter: A forced alignment system based on deep neural networks and interpolation
di: Kelley, Matthew C., et al.
Pubblicazione: (2023)
di: Kelley, Matthew C., et al.
Pubblicazione: (2023)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
di: Grossman, Raymond, et al.
Pubblicazione: (2025)
di: Grossman, Raymond, et al.
Pubblicazione: (2025)
Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
di: Wang, Ke, et al.
Pubblicazione: (2025)
di: Wang, Ke, et al.
Pubblicazione: (2025)
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
di: Wang, Bin, et al.
Pubblicazione: (2025)
di: Wang, Bin, et al.
Pubblicazione: (2025)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
di: Cui, Wenqian, et al.
Pubblicazione: (2026)
di: Cui, Wenqian, et al.
Pubblicazione: (2026)
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
di: Peri, Raghuveer, et al.
Pubblicazione: (2024)
di: Peri, Raghuveer, et al.
Pubblicazione: (2024)
SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
di: Pozorski, Paweł, et al.
Pubblicazione: (2026)
di: Pozorski, Paweł, et al.
Pubblicazione: (2026)
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment
di: Wang, Ke, et al.
Pubblicazione: (2025)
di: Wang, Ke, et al.
Pubblicazione: (2025)
The ART of Conversation: Measuring Phonetic Convergence and Deliberate Imitation in L2-Speech with a Siamese RNN
di: Yuan, Zheng, et al.
Pubblicazione: (2023)
di: Yuan, Zheng, et al.
Pubblicazione: (2023)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
di: Yang, Hao, et al.
Pubblicazione: (2024)
di: Yang, Hao, et al.
Pubblicazione: (2024)
What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models
di: Çoban, Enis Berk, et al.
Pubblicazione: (2024)
di: Çoban, Enis Berk, et al.
Pubblicazione: (2024)
CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
di: Wu, Shangda, et al.
Pubblicazione: (2024)
di: Wu, Shangda, et al.
Pubblicazione: (2024)
Multimodal Magic Elevating Depression Detection with a Fusion of Text and Audio Intelligence
di: Gan, Lindy, et al.
Pubblicazione: (2025)
di: Gan, Lindy, et al.
Pubblicazione: (2025)
High-Fidelity Neural Phonetic Posteriorgrams
di: Churchwell, Cameron, et al.
Pubblicazione: (2024)
di: Churchwell, Cameron, et al.
Pubblicazione: (2024)
EMMeTT: Efficient Multimodal Machine Translation Training
di: Żelasko, Piotr, et al.
Pubblicazione: (2024)
di: Żelasko, Piotr, et al.
Pubblicazione: (2024)
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
di: Hu, Jiliang, et al.
Pubblicazione: (2024)
di: Hu, Jiliang, et al.
Pubblicazione: (2024)
OCR-Enhanced Multimodal ASR Can Read While Listening
di: Chen, Junli, et al.
Pubblicazione: (2026)
di: Chen, Junli, et al.
Pubblicazione: (2026)
How to Learn a New Language? An Efficient Solution for Self-Supervised Learning Models Unseen Languages Adaption in Low-Resource Scenario
di: Wang, Shih-Heng, et al.
Pubblicazione: (2024)
di: Wang, Shih-Heng, et al.
Pubblicazione: (2024)
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
di: Zhou, Xuanru, et al.
Pubblicazione: (2025)
di: Zhou, Xuanru, et al.
Pubblicazione: (2025)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
Phonetic Richness for Improved Automatic Speaker Verification
di: Klein, Nicholas, et al.
Pubblicazione: (2024)
di: Klein, Nicholas, et al.
Pubblicazione: (2024)
PhiNet: Speaker Verification with Phonetic Interpretability
di: Ma, Yi, et al.
Pubblicazione: (2026)
di: Ma, Yi, et al.
Pubblicazione: (2026)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
di: Wang, Xuechen, et al.
Pubblicazione: (2024)
di: Wang, Xuechen, et al.
Pubblicazione: (2024)
Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation
di: Lei, Ligong, et al.
Pubblicazione: (2026)
di: Lei, Ligong, et al.
Pubblicazione: (2026)
GatedxLSTM: A Multimodal Affective Computing Approach for Emotion Recognition in Conversations
di: Li, Yupei, et al.
Pubblicazione: (2025)
di: Li, Yupei, et al.
Pubblicazione: (2025)
What Do Speech Foundation Models Not Learn About Speech?
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Phonetic Segmentation of the UCLA Phonetics Lab Archive
di: Chodroff, Eleanor, et al.
Pubblicazione: (2024) -
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
di: Zhou, Kun, et al.
Pubblicazione: (2024) -
Phonetic Error Analysis of Raw Waveform Acoustic Models with Parametric and Non-Parametric CNNs
di: Loweimi, Erfan, et al.
Pubblicazione: (2024) -
A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs
di: Šegedin, Bruno Ferenc
Pubblicazione: (2025) -
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
di: Yusuyin, Saierdaer, et al.
Pubblicazione: (2024)