Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Rousso, Rotem, Cohen, Eyal, Keshet, Joseph, Chodroff, Eleanor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
by: Cohen, Eyal, et al.
Published: (2025)
by: Cohen, Eyal, et al.
Published: (2025)
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
by: Farhadipour, Aref, et al.
Published: (2026)
by: Farhadipour, Aref, et al.
Published: (2026)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
by: Dissen, Yehoshua, et al.
Published: (2024)
by: Dissen, Yehoshua, et al.
Published: (2024)
How Does a Deep Neural Network Look at Lexical Stress in English Words?
by: Allouche, Itai, et al.
Published: (2025)
by: Allouche, Itai, et al.
Published: (2025)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
by: Chodroff, Eleanor, et al.
Published: (2024)
by: Chodroff, Eleanor, et al.
Published: (2024)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
by: Zhang, Miao, et al.
Published: (2025)
by: Zhang, Miao, et al.
Published: (2025)
ZIPA: A family of efficient models for multilingual phone recognition
by: Zhu, Jian, et al.
Published: (2025)
by: Zhu, Jian, et al.
Published: (2025)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
by: Segal-Feldman, Yael, et al.
Published: (2024)
by: Segal-Feldman, Yael, et al.
Published: (2024)
TidyVoice 2026 Challenge Evaluation Plan
by: Farhadipour, Aref, et al.
Published: (2026)
by: Farhadipour, Aref, et al.
Published: (2026)
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
by: Benita, Roi, et al.
Published: (2023)
by: Benita, Roi, et al.
Published: (2023)
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
by: Li, Li, et al.
Published: (2026)
by: Li, Li, et al.
Published: (2026)
Comparative Analysis of ASR Methods for Speech Deepfake Detection
by: Salvi, Davide, et al.
Published: (2024)
by: Salvi, Davide, et al.
Published: (2024)
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
by: Raissi, Tina, et al.
Published: (2024)
by: Raissi, Tina, et al.
Published: (2024)
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
by: Yen, Hao, et al.
Published: (2026)
by: Yen, Hao, et al.
Published: (2026)
Beyond Transcription: Mechanistic Interpretability in ASR
by: Glazer, Neta, et al.
Published: (2025)
by: Glazer, Neta, et al.
Published: (2025)
PatchDSU: Uncertainty Modeling for Out of Distribution Generalization in Keyword Spotting
by: Chernyak, Bronya Roni, et al.
Published: (2025)
by: Chernyak, Bronya Roni, et al.
Published: (2025)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
by: Moriya, Takafumi, et al.
Published: (2025)
by: Moriya, Takafumi, et al.
Published: (2025)
Algorithms for Collaborative Harmonization
by: Briman, Eyal, et al.
Published: (2025)
by: Briman, Eyal, et al.
Published: (2025)
Inverse-Hessian Regularization for Continual Learning in ASR
by: Eeckt, Steven Vander, et al.
Published: (2026)
by: Eeckt, Steven Vander, et al.
Published: (2026)
BFA: Real-time Multilingual Text-to-speech Forced Alignment
by: Rehman, Abdul, et al.
Published: (2025)
by: Rehman, Abdul, et al.
Published: (2025)
A Low-Complexity Speech Codec Using Parametric Dithering for ASR
by: Murray, Ellison, et al.
Published: (2025)
by: Murray, Ellison, et al.
Published: (2025)
SOT Triggered Neural Clustering for Speaker Attributed ASR
by: Zheng, Xianrui, et al.
Published: (2024)
by: Zheng, Xianrui, et al.
Published: (2024)
An investigation of modularity for noise robustness in conformer-based ASR
by: de Gibson, Louise Coppieters, et al.
Published: (2024)
by: de Gibson, Louise Coppieters, et al.
Published: (2024)
DNCASR: End-to-End Training for Speaker-Attributed ASR
by: Zheng, Xianrui, et al.
Published: (2025)
by: Zheng, Xianrui, et al.
Published: (2025)
HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
by: Turetzky, Arnon, et al.
Published: (2024)
by: Turetzky, Arnon, et al.
Published: (2024)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
by: Kumar, Shashi, et al.
Published: (2024)
by: Kumar, Shashi, et al.
Published: (2024)
Towards a Single ASR Model That Generalizes to Disordered Speech
by: Tobin, Jimmy, et al.
Published: (2024)
by: Tobin, Jimmy, et al.
Published: (2024)
NLE: Non-autoregressive LLM-based ASR by Transcript Editing
by: Dekel, Avihu, et al.
Published: (2026)
by: Dekel, Avihu, et al.
Published: (2026)
Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time
by: Allouche, Itai, et al.
Published: (2026)
by: Allouche, Itai, et al.
Published: (2026)
CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
by: Shankar, Natarajan Balaji, et al.
Published: (2025)
by: Shankar, Natarajan Balaji, et al.
Published: (2025)
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
by: Shi, Mohan, et al.
Published: (2025)
by: Shi, Mohan, et al.
Published: (2025)
Target Speaker ASR with Whisper
by: Polok, Alexander, et al.
Published: (2024)
by: Polok, Alexander, et al.
Published: (2024)
Index-ASR Technical Report
by: Song, Zheshu, et al.
Published: (2025)
by: Song, Zheshu, et al.
Published: (2025)
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
by: Fujita, Yuya, et al.
Published: (2024)
by: Fujita, Yuya, et al.
Published: (2024)
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
by: Wu, Ke, et al.
Published: (2026)
by: Wu, Ke, et al.
Published: (2026)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
by: Tsai, Kai-Chen, et al.
Published: (2026)
by: Tsai, Kai-Chen, et al.
Published: (2026)
ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy
by: Wu, Ya-Tse, et al.
Published: (2026)
by: Wu, Ya-Tse, et al.
Published: (2026)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
by: Kim, Jongsuk, et al.
Published: (2025)
by: Kim, Jongsuk, et al.
Published: (2025)
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
by: Saon, George, et al.
Published: (2026)
by: Saon, George, et al.
Published: (2026)
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
by: Joshi, Sakshi, et al.
Published: (2025)
by: Joshi, Sakshi, et al.
Published: (2025)
Similar Items
-
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
by: Cohen, Eyal, et al.
Published: (2025) -
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
by: Farhadipour, Aref, et al.
Published: (2026) -
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
by: Dissen, Yehoshua, et al.
Published: (2024) -
How Does a Deep Neural Network Look at Lexical Stress in English Words?
by: Allouche, Itai, et al.
Published: (2025) -
Phonetic Segmentation of the UCLA Phonetics Lab Archive
by: Chodroff, Eleanor, et al.
Published: (2024)