Saved in:
| Main Authors: | Kelley, Matthew C., Perry, Scott James, Tucker, Benjamin V. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2310.15425 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gradient boundaries through confidence intervals for forced alignment estimates using model ensembles
by: Kelley, Matthew C.
Published: (2025)
by: Kelley, Matthew C.
Published: (2025)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
by: Chodroff, Eleanor, et al.
Published: (2024)
by: Chodroff, Eleanor, et al.
Published: (2024)
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
by: Yamashita, Natsuo, et al.
Published: (2025)
by: Yamashita, Natsuo, et al.
Published: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
by: Zhou, Kun, et al.
Published: (2024)
by: Zhou, Kun, et al.
Published: (2024)
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
by: Zhu, Jian, et al.
Published: (2023)
by: Zhu, Jian, et al.
Published: (2023)
Multimodal Input Aids a Bayesian Model of Phonetic Learning
by: Zhi, Sophia, et al.
Published: (2024)
by: Zhi, Sophia, et al.
Published: (2024)
A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs
by: Šegedin, Bruno Ferenc
Published: (2025)
by: Šegedin, Bruno Ferenc
Published: (2025)
Phonetic Error Analysis of Raw Waveform Acoustic Models with Parametric and Non-Parametric CNNs
by: Loweimi, Erfan, et al.
Published: (2024)
by: Loweimi, Erfan, et al.
Published: (2024)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
by: Yusuyin, Saierdaer, et al.
Published: (2024)
by: Yusuyin, Saierdaer, et al.
Published: (2024)
PAST: Phonetic-Acoustic Speech Tokenizer
by: Har-Tuv, Nadav, et al.
Published: (2025)
by: Har-Tuv, Nadav, et al.
Published: (2025)
Basic syntax from speech: Spontaneous concatenation in unsupervised deep neural networks
by: Beguš, Gašper, et al.
Published: (2023)
by: Beguš, Gašper, et al.
Published: (2023)
(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test
by: Bleeck, Stefan
Published: (2025)
by: Bleeck, Stefan
Published: (2025)
High-Fidelity Neural Phonetic Posteriorgrams
by: Churchwell, Cameron, et al.
Published: (2024)
by: Churchwell, Cameron, et al.
Published: (2024)
Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
ISPA: Inter-Species Phonetic Alphabet for Transcribing Animal Sounds
by: Hagiwara, Masato, et al.
Published: (2024)
by: Hagiwara, Masato, et al.
Published: (2024)
PhiNet: Speaker Verification with Phonetic Interpretability
by: Ma, Yi, et al.
Published: (2026)
by: Ma, Yi, et al.
Published: (2026)
Phonetic Richness for Improved Automatic Speaker Verification
by: Klein, Nicholas, et al.
Published: (2024)
by: Klein, Nicholas, et al.
Published: (2024)
Revealing the Hidden Temporal Structure of HubertSoft Embeddings based on the Russian Phonetic Corpus
by: Ananeva, Anastasia, et al.
Published: (2025)
by: Ananeva, Anastasia, et al.
Published: (2025)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
by: Li, Xingyuan, et al.
Published: (2024)
by: Li, Xingyuan, et al.
Published: (2024)
Transferable speech-to-text large language model alignment module
by: Wu, Boyong, et al.
Published: (2024)
by: Wu, Boyong, et al.
Published: (2024)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
by: Yamashita, Natsuo, et al.
Published: (2026)
by: Yamashita, Natsuo, et al.
Published: (2026)
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
by: Choi, Kwanghee, et al.
Published: (2026)
by: Choi, Kwanghee, et al.
Published: (2026)
The ART of Conversation: Measuring Phonetic Convergence and Deliberate Imitation in L2-Speech with a Siamese RNN
by: Yuan, Zheng, et al.
Published: (2023)
by: Yuan, Zheng, et al.
Published: (2023)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
by: Zhang, Miao, et al.
Published: (2025)
by: Zhang, Miao, et al.
Published: (2025)
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
by: Li, Jialu, et al.
Published: (2023)
by: Li, Jialu, et al.
Published: (2023)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
by: Han, Seungu, et al.
Published: (2026)
by: Han, Seungu, et al.
Published: (2026)
A circular microphone array with virtual microphones based on acoustics-informed neural networks
by: Zhao, Sipei, et al.
Published: (2024)
by: Zhao, Sipei, et al.
Published: (2024)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
by: Kloots, Marianne de Heer, et al.
Published: (2024)
by: Kloots, Marianne de Heer, et al.
Published: (2024)
SPO-CLAPScore: Enhancing CLAP-based alignment prediction system with Standardize Preference Optimization, for the first XACLE Challenge
by: Takano, Taisei, et al.
Published: (2026)
by: Takano, Taisei, et al.
Published: (2026)
PSST! Prosodic Speech Segmentation with Transformers
by: Roll, Nathan, et al.
Published: (2023)
by: Roll, Nathan, et al.
Published: (2023)
SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
by: Pozorski, Paweł, et al.
Published: (2026)
by: Pozorski, Paweł, et al.
Published: (2026)
An adaptive filter bank based neural network approach for time delay estimation and speech enhancement
by: Ma, Lu
Published: (2025)
by: Ma, Lu
Published: (2025)
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network
by: Ma, Wenbo, et al.
Published: (2024)
by: Ma, Wenbo, et al.
Published: (2024)
Lightweight Audio Segmentation for Long-form Speech Translation
by: Lee, Jaesong, et al.
Published: (2024)
by: Lee, Jaesong, et al.
Published: (2024)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
by: Sanders, Nicholas, et al.
Published: (2025)
by: Sanders, Nicholas, et al.
Published: (2025)
Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
by: Shukla, Sakshi Deo, et al.
Published: (2024)
by: Shukla, Sakshi Deo, et al.
Published: (2024)
A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
by: Wang, Xiaopeng, et al.
Published: (2024)
by: Wang, Xiaopeng, et al.
Published: (2024)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
by: Liu, Oli Danyi, et al.
Published: (2024)
by: Liu, Oli Danyi, et al.
Published: (2024)
REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASR
by: Tseng, Liang-Hsuan, et al.
Published: (2024)
by: Tseng, Liang-Hsuan, et al.
Published: (2024)
Similar Items
-
Gradient boundaries through confidence intervals for forced alignment estimates using model ensembles
by: Kelley, Matthew C.
Published: (2025) -
Phonetic Segmentation of the UCLA Phonetics Lab Archive
by: Chodroff, Eleanor, et al.
Published: (2024) -
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
by: Yamashita, Natsuo, et al.
Published: (2025) -
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
by: Zhou, Kun, et al.
Published: (2024) -
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
by: Zhu, Jian, et al.
Published: (2023)