Gradient boundaries through confidence intervals for forced alignment estimates using model ensembles
Fuente:
arXiv
Guardado en:
| Autor principal: | Kelley, Matthew C. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Mason-Alberta Phonetic Segmenter: A forced alignment system based on deep neural networks and interpolation
por: Kelley, Matthew C., et al.
Publicado: (2023)
por: Kelley, Matthew C., et al.
Publicado: (2023)
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
por: Zhu, Jian, et al.
Publicado: (2023)
por: Zhu, Jian, et al.
Publicado: (2023)
Text-only adaptation in LLM-based ASR through text denoising
por: Carofilis, Andrés, et al.
Publicado: (2026)
por: Carofilis, Andrés, et al.
Publicado: (2026)
Digits micro-model for accurate and secure transactions
por: Chhablani, Chirag, et al.
Publicado: (2024)
por: Chhablani, Chirag, et al.
Publicado: (2024)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
por: Anand, Srija, et al.
Publicado: (2024)
por: Anand, Srija, et al.
Publicado: (2024)
Efficient VoIP Communications through LLM-based Real-Time Speech Reconstruction and Call Prioritization for Emergency Services
por: Venkateshperumal, Danush, et al.
Publicado: (2024)
por: Venkateshperumal, Danush, et al.
Publicado: (2024)
Transferable speech-to-text large language model alignment module
por: Wu, Boyong, et al.
Publicado: (2024)
por: Wu, Boyong, et al.
Publicado: (2024)
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
por: Lei, Zhihong, et al.
Publicado: (2024)
por: Lei, Zhihong, et al.
Publicado: (2024)
Modeling Overlapped Speech with Shuffles
por: Wiesner, Matthew, et al.
Publicado: (2026)
por: Wiesner, Matthew, et al.
Publicado: (2026)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
por: Li, Yingting, et al.
Publicado: (2024)
por: Li, Yingting, et al.
Publicado: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
por: Li, Xingyuan, et al.
Publicado: (2024)
por: Li, Xingyuan, et al.
Publicado: (2024)
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving
por: Shankar, Bhavani, et al.
Publicado: (2024)
por: Shankar, Bhavani, et al.
Publicado: (2024)
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
por: You, Jian, et al.
Publicado: (2024)
por: You, Jian, et al.
Publicado: (2024)
How Does a Deep Neural Network Look at Lexical Stress in English Words?
por: Allouche, Itai, et al.
Publicado: (2025)
por: Allouche, Itai, et al.
Publicado: (2025)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
por: Bataev, Vladimir, et al.
Publicado: (2023)
por: Bataev, Vladimir, et al.
Publicado: (2023)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
por: Fujita, Kenichi, et al.
Publicado: (2024)
por: Fujita, Kenichi, et al.
Publicado: (2024)
Tempo estimation as fully self-supervised binary classification
por: Henkel, Florian, et al.
Publicado: (2024)
por: Henkel, Florian, et al.
Publicado: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
por: Park, Taejin, et al.
Publicado: (2024)
por: Park, Taejin, et al.
Publicado: (2024)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
por: Puvvada, Krishna C., et al.
Publicado: (2024)
por: Puvvada, Krishna C., et al.
Publicado: (2024)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
por: Ngo, Huong, et al.
Publicado: (2025)
por: Ngo, Huong, et al.
Publicado: (2025)
Bayesian Low-Rank Factorization for Robust Model Adaptation
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
Voice Impression Control in Zero-Shot TTS
por: Fujita, Kenichi, et al.
Publicado: (2025)
por: Fujita, Kenichi, et al.
Publicado: (2025)
WhisperRT -- Turning Whisper into a Causal Streaming Model
por: Krichli, Tomer, et al.
Publicado: (2025)
por: Krichli, Tomer, et al.
Publicado: (2025)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
por: Amooie, Reihaneh, et al.
Publicado: (2025)
por: Amooie, Reihaneh, et al.
Publicado: (2025)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
por: Eren, Eray, et al.
Publicado: (2025)
por: Eren, Eray, et al.
Publicado: (2025)
Label-Context-Dependent Internal Language Model Estimation for CTC
por: Yang, Zijian, et al.
Publicado: (2025)
por: Yang, Zijian, et al.
Publicado: (2025)
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
por: Ghosh, Sreyan, et al.
Publicado: (2025)
por: Ghosh, Sreyan, et al.
Publicado: (2025)
TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics
por: Lin, Yi-Cheng, et al.
Publicado: (2025)
por: Lin, Yi-Cheng, et al.
Publicado: (2025)
Error Analysis in a Modular Meeting Transcription System
por: Vieting, Peter, et al.
Publicado: (2025)
por: Vieting, Peter, et al.
Publicado: (2025)
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
por: Chang, Sungkyun, et al.
Publicado: (2025)
por: Chang, Sungkyun, et al.
Publicado: (2025)
Unified Learnable 2D Convolutional Feature Extraction for ASR
por: Vieting, Peter, et al.
Publicado: (2025)
por: Vieting, Peter, et al.
Publicado: (2025)
Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
por: Falai, Alessio, et al.
Publicado: (2025)
por: Falai, Alessio, et al.
Publicado: (2025)
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
por: Beyene, Luel Hagos, et al.
Publicado: (2025)
por: Beyene, Luel Hagos, et al.
Publicado: (2025)
Breathing and Semantic Pause Detection and Exertion-Level Classification in Post-Exercise Speech
por: Wang, Yuyu, et al.
Publicado: (2025)
por: Wang, Yuyu, et al.
Publicado: (2025)
SpeakStream: Streaming Text-to-Speech with Interleaved Data
por: Bai, Richard He, et al.
Publicado: (2025)
por: Bai, Richard He, et al.
Publicado: (2025)
The State Of TTS: A Case Study with Human Fooling Rates
por: Varadhan, Praveen Srinivasa, et al.
Publicado: (2025)
por: Varadhan, Praveen Srinivasa, et al.
Publicado: (2025)
Using Phonemes in cascaded S2S translation pipeline
por: Pilz, Rene, et al.
Publicado: (2025)
por: Pilz, Rene, et al.
Publicado: (2025)
WEE-Therapy: A Mixture of Weak Encoders Framework for Psychological Counseling Dialogue Analysis
por: Kang, Yongqi, et al.
Publicado: (2025)
por: Kang, Yongqi, et al.
Publicado: (2025)
Text to Speech System for Meitei Mayek Script
por: Irengbam, Gangular Singh, et al.
Publicado: (2025)
por: Irengbam, Gangular Singh, et al.
Publicado: (2025)
Ejemplares similares
-
The Mason-Alberta Phonetic Segmenter: A forced alignment system based on deep neural networks and interpolation
por: Kelley, Matthew C., et al.
Publicado: (2023) -
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
por: Zhu, Jian, et al.
Publicado: (2023) -
Text-only adaptation in LLM-based ASR through text denoising
por: Carofilis, Andrés, et al.
Publicado: (2026) -
Digits micro-model for accurate and secure transactions
por: Chhablani, Chirag, et al.
Publicado: (2024) -
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
por: Anand, Srija, et al.
Publicado: (2024)