The Mason-Alberta Phonetic Segmenter: A forced alignment system based on deep neural networks and interpolation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kelley, Matthew C., Perry, Scott James, Tucker, Benjamin V. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
Gradient boundaries through confidence intervals for forced alignment estimates using model ensembles
von: Kelley, Matthew C.
Veröffentlicht: (2025)
von: Kelley, Matthew C.
Veröffentlicht: (2025)
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2025)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
von: Zhu, Jian, et al.
Veröffentlicht: (2023)
von: Zhu, Jian, et al.
Veröffentlicht: (2023)
Multimodal Input Aids a Bayesian Model of Phonetic Learning
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs
von: Šegedin, Bruno Ferenc
Veröffentlicht: (2025)
von: Šegedin, Bruno Ferenc
Veröffentlicht: (2025)
Phonetic Error Analysis of Raw Waveform Acoustic Models with Parametric and Non-Parametric CNNs
von: Loweimi, Erfan, et al.
Veröffentlicht: (2024)
von: Loweimi, Erfan, et al.
Veröffentlicht: (2024)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
High-Fidelity Neural Phonetic Posteriorgrams
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test
von: Bleeck, Stefan
Veröffentlicht: (2025)
von: Bleeck, Stefan
Veröffentlicht: (2025)
PhiNet: Speaker Verification with Phonetic Interpretability
von: Ma, Yi, et al.
Veröffentlicht: (2026)
von: Ma, Yi, et al.
Veröffentlicht: (2026)
Phonetic Richness for Improved Automatic Speaker Verification
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
Basic syntax from speech: Spontaneous concatenation in unsupervised deep neural networks
von: Beguš, Gašper, et al.
Veröffentlicht: (2023)
von: Beguš, Gašper, et al.
Veröffentlicht: (2023)
Revealing the Hidden Temporal Structure of HubertSoft Embeddings based on the Russian Phonetic Corpus
von: Ananeva, Anastasia, et al.
Veröffentlicht: (2025)
von: Ananeva, Anastasia, et al.
Veröffentlicht: (2025)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2026)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2026)
Transferable speech-to-text large language model alignment module
von: Wu, Boyong, et al.
Veröffentlicht: (2024)
von: Wu, Boyong, et al.
Veröffentlicht: (2024)
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation
von: Roh, Jaechul, et al.
Veröffentlicht: (2025)
von: Roh, Jaechul, et al.
Veröffentlicht: (2025)
ISPA: Inter-Species Phonetic Alphabet for Transcribing Animal Sounds
von: Hagiwara, Masato, et al.
Veröffentlicht: (2024)
von: Hagiwara, Masato, et al.
Veröffentlicht: (2024)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
von: Li, Jialu, et al.
Veröffentlicht: (2023)
von: Li, Jialu, et al.
Veröffentlicht: (2023)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026)
von: Han, Seungu, et al.
Veröffentlicht: (2026)
A circular microphone array with virtual microphones based on acoustics-informed neural networks
von: Zhao, Sipei, et al.
Veröffentlicht: (2024)
von: Zhao, Sipei, et al.
Veröffentlicht: (2024)
SPO-CLAPScore: Enhancing CLAP-based alignment prediction system with Standardize Preference Optimization, for the first XACLE Challenge
von: Takano, Taisei, et al.
Veröffentlicht: (2026)
von: Takano, Taisei, et al.
Veröffentlicht: (2026)
An adaptive filter bank based neural network approach for time delay estimation and speech enhancement
von: Ma, Lu
Veröffentlicht: (2025)
von: Ma, Lu
Veröffentlicht: (2025)
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network
von: Ma, Wenbo, et al.
Veröffentlicht: (2024)
von: Ma, Wenbo, et al.
Veröffentlicht: (2024)
SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
von: Pozorski, Paweł, et al.
Veröffentlicht: (2026)
von: Pozorski, Paweł, et al.
Veröffentlicht: (2026)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
The ART of Conversation: Measuring Phonetic Convergence and Deliberate Imitation in L2-Speech with a Siamese RNN
von: Yuan, Zheng, et al.
Veröffentlicht: (2023)
von: Yuan, Zheng, et al.
Veröffentlicht: (2023)
PSST! Prosodic Speech Segmentation with Transformers
von: Roll, Nathan, et al.
Veröffentlicht: (2023)
von: Roll, Nathan, et al.
Veröffentlicht: (2023)
Towards detecting the pathological subharmonic voicing with fully convolutional neural networks
von: Ikuma, Takeshi, et al.
Veröffentlicht: (2025)
von: Ikuma, Takeshi, et al.
Veröffentlicht: (2025)
Lightweight Audio Segmentation for Long-form Speech Translation
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
Multispecies bird sound recognition using a fully convolutional neural network
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
von: Shukla, Sakshi Deo, et al.
Veröffentlicht: (2024)
von: Shukla, Sakshi Deo, et al.
Veröffentlicht: (2024)
A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
von: Wang, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiaopeng, et al.
Veröffentlicht: (2024)
REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASR
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2024)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024) -
Gradient boundaries through confidence intervals for forced alignment estimates using model ensembles
von: Kelley, Matthew C.
Veröffentlicht: (2025) -
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2025) -
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024) -
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
von: Zhu, Jian, et al.
Veröffentlicht: (2023)