A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs
Fuente:
arXiv
Saved in:
| Main Author: | Šegedin, Bruno Ferenc |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Phonetic Error Analysis of Raw Waveform Acoustic Models with Parametric and Non-Parametric CNNs
by: Loweimi, Erfan, et al.
Published: (2024)
by: Loweimi, Erfan, et al.
Published: (2024)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
by: Chodroff, Eleanor, et al.
Published: (2024)
by: Chodroff, Eleanor, et al.
Published: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
by: Li, Xingyuan, et al.
Published: (2024)
by: Li, Xingyuan, et al.
Published: (2024)
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
by: Yamashita, Natsuo, et al.
Published: (2025)
by: Yamashita, Natsuo, et al.
Published: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
by: Zhou, Kun, et al.
Published: (2024)
by: Zhou, Kun, et al.
Published: (2024)
Multimodal Input Aids a Bayesian Model of Phonetic Learning
by: Zhi, Sophia, et al.
Published: (2024)
by: Zhi, Sophia, et al.
Published: (2024)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
by: Yusuyin, Saierdaer, et al.
Published: (2024)
by: Yusuyin, Saierdaer, et al.
Published: (2024)
On the Contribution of Lexical Features to Speech Emotion Recognition
by: Combei, David
Published: (2025)
by: Combei, David
Published: (2025)
(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test
by: Bleeck, Stefan
Published: (2025)
by: Bleeck, Stefan
Published: (2025)
Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
by: Har-Tuv, Nadav, et al.
Published: (2025)
by: Har-Tuv, Nadav, et al.
Published: (2025)
Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech
by: Räsänen, Okko, et al.
Published: (2024)
by: Räsänen, Okko, et al.
Published: (2024)
ISPA: Inter-Species Phonetic Alphabet for Transcribing Animal Sounds
by: Hagiwara, Masato, et al.
Published: (2024)
by: Hagiwara, Masato, et al.
Published: (2024)
Proactive Hearing Assistants that Isolate Egocentric Conversations
by: Hu, Guilin, et al.
Published: (2025)
by: Hu, Guilin, et al.
Published: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
by: Park, Chanho, et al.
Published: (2024)
by: Park, Chanho, et al.
Published: (2024)
The Mason-Alberta Phonetic Segmenter: A forced alignment system based on deep neural networks and interpolation
by: Kelley, Matthew C., et al.
Published: (2023)
by: Kelley, Matthew C., et al.
Published: (2023)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
by: Choi, Kwanghee, et al.
Published: (2026)
by: Choi, Kwanghee, et al.
Published: (2026)
PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques
by: Lin, Tzu-Han, et al.
Published: (2024)
by: Lin, Tzu-Han, et al.
Published: (2024)
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
The ART of Conversation: Measuring Phonetic Convergence and Deliberate Imitation in L2-Speech with a Siamese RNN
by: Yuan, Zheng, et al.
Published: (2023)
by: Yuan, Zheng, et al.
Published: (2023)
The Voice of Equity: A Systematic Evaluation of Bias Mitigation Techniques for Speech-Based Cognitive Impairment Detection Across Architectures and Demographics
by: Haghbin, Yasaman, et al.
Published: (2026)
by: Haghbin, Yasaman, et al.
Published: (2026)
High-Fidelity Neural Phonetic Posteriorgrams
by: Churchwell, Cameron, et al.
Published: (2024)
by: Churchwell, Cameron, et al.
Published: (2024)
PhiNet: Speaker Verification with Phonetic Interpretability
by: Ma, Yi, et al.
Published: (2026)
by: Ma, Yi, et al.
Published: (2026)
Phonetic Richness for Improved Automatic Speaker Verification
by: Klein, Nicholas, et al.
Published: (2024)
by: Klein, Nicholas, et al.
Published: (2024)
Exploring Dynamic Parameters for Vietnamese Gender-Independent ASR
by: Leang, Sotheara, et al.
Published: (2025)
by: Leang, Sotheara, et al.
Published: (2025)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
by: Kloots, Marianne de Heer, et al.
Published: (2024)
by: Kloots, Marianne de Heer, et al.
Published: (2024)
How Does a Deep Neural Network Look at Lexical Stress in English Words?
by: Allouche, Itai, et al.
Published: (2025)
by: Allouche, Itai, et al.
Published: (2025)
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
by: Lou, Haowei, et al.
Published: (2025)
by: Lou, Haowei, et al.
Published: (2025)
UniCoM: A Universal Code-Switching Speech Generator
by: Lee, Sangmin, et al.
Published: (2025)
by: Lee, Sangmin, et al.
Published: (2025)
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
by: He, Haorui, et al.
Published: (2025)
by: He, Haorui, et al.
Published: (2025)
An experiment on an automated literature survey of data-driven speech enhancement methods
by: Santos, Arthur dos, et al.
Published: (2023)
by: Santos, Arthur dos, et al.
Published: (2023)
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
by: Song, Yaodong, et al.
Published: (2025)
by: Song, Yaodong, et al.
Published: (2025)
Braille-to-Speech Generator: Audio Generation Based on Joint Fine-Tuning of CLIP and Fastspeech2
by: Xu, Chun, et al.
Published: (2024)
by: Xu, Chun, et al.
Published: (2024)
WavMark: Watermarking for Audio Generation
by: Chen, Guangyu, et al.
Published: (2023)
by: Chen, Guangyu, et al.
Published: (2023)
Generative Expressive Conversational Speech Synthesis
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Dialectal Coverage And Generalization in Arabic Speech Recognition
by: Djanibekov, Amirbek, et al.
Published: (2024)
by: Djanibekov, Amirbek, et al.
Published: (2024)
Cross-Utterance Conditioned VAE for Speech Generation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
by: Wang, Hankun, et al.
Published: (2024)
by: Wang, Hankun, et al.
Published: (2024)
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
by: Lu, Ke-Han, et al.
Published: (2025)
by: Lu, Ke-Han, et al.
Published: (2025)
Versatile Framework for Song Generation with Prompt-based Control
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Similar Items
-
Phonetic Error Analysis of Raw Waveform Acoustic Models with Parametric and Non-Parametric CNNs
by: Loweimi, Erfan, et al.
Published: (2024) -
Phonetic Segmentation of the UCLA Phonetics Lab Archive
by: Chodroff, Eleanor, et al.
Published: (2024) -
Phonetic and Lexical Discovery of a Canine Language using HuBERT
by: Li, Xingyuan, et al.
Published: (2024) -
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
by: Yamashita, Natsuo, et al.
Published: (2025) -
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
by: Zhou, Kun, et al.
Published: (2024)