Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Kevin, Chkhetiani, Luka, Ramirez, Francis McCann, Khare, Yash, Vanzo, Andrea, Liang, Michael, Martin, Sergio Ramirez, Oexle, Gabriel, Bousbib, Ruben, Peyash, Taufiquzzaman, Nguyen, Michael, Pulliam, Dillon, Donato, Domenic |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Anatomy of Industrial Scale Multilingual ASR
di: Ramirez, Francis McCann, et al.
Pubblicazione: (2024)
di: Ramirez, Francis McCann, et al.
Pubblicazione: (2024)
Universal-2-TF: Robust All-Neural Text Formatting for ASR
di: Khare, Yash, et al.
Pubblicazione: (2025)
di: Khare, Yash, et al.
Pubblicazione: (2025)
An open-source voice type classifier for child-centered daylong recordings
di: Lavechin, Marvin, et al.
Pubblicazione: (2020)
di: Lavechin, Marvin, et al.
Pubblicazione: (2020)
Promptformer: Prompted Conformer Transducer for ASR
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation
di: Cai, Danwei, et al.
Pubblicazione: (2023)
di: Cai, Danwei, et al.
Pubblicazione: (2023)
Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
di: Wang, Tianzi, et al.
Pubblicazione: (2025)
di: Wang, Tianzi, et al.
Pubblicazione: (2025)
Cross-utterance ASR Rescoring with Graph-based Label Propagation
di: Tankasala, Srinath, et al.
Pubblicazione: (2023)
di: Tankasala, Srinath, et al.
Pubblicazione: (2023)
Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
di: Li, Zhaoqing, et al.
Pubblicazione: (2025)
di: Li, Zhaoqing, et al.
Pubblicazione: (2025)
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
di: Li, Li, et al.
Pubblicazione: (2026)
di: Li, Li, et al.
Pubblicazione: (2026)
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
di: Lee, Jaeyoung, et al.
Pubblicazione: (2026)
di: Lee, Jaeyoung, et al.
Pubblicazione: (2026)
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
di: Yen, Hao, et al.
Pubblicazione: (2026)
di: Yen, Hao, et al.
Pubblicazione: (2026)
DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR
di: Huybrechts, Goeric, et al.
Pubblicazione: (2023)
di: Huybrechts, Goeric, et al.
Pubblicazione: (2023)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
di: Moriya, Takafumi, et al.
Pubblicazione: (2025)
di: Moriya, Takafumi, et al.
Pubblicazione: (2025)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
PromptASR for contextualized ASR with controllable style
di: Yang, Xiaoyu, et al.
Pubblicazione: (2023)
di: Yang, Xiaoyu, et al.
Pubblicazione: (2023)
ASR Benchmarking: Need for a More Representative Conversational Dataset
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2024)
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2024)
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
di: Ye, Shuaishuai, et al.
Pubblicazione: (2024)
di: Ye, Shuaishuai, et al.
Pubblicazione: (2024)
Target Speaker ASR with Whisper
di: Polok, Alexander, et al.
Pubblicazione: (2024)
di: Polok, Alexander, et al.
Pubblicazione: (2024)
Index-ASR Technical Report
di: Song, Zheshu, et al.
Pubblicazione: (2025)
di: Song, Zheshu, et al.
Pubblicazione: (2025)
Inverse-Hessian Regularization for Continual Learning in ASR
di: Eeckt, Steven Vander, et al.
Pubblicazione: (2026)
di: Eeckt, Steven Vander, et al.
Pubblicazione: (2026)
Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
di: Yan, Brian, et al.
Pubblicazione: (2024)
di: Yan, Brian, et al.
Pubblicazione: (2024)
SSCFormer: Push the Limit of Chunk-wise Conformer for Streaming ASR Using Sequentially Sampled Chunks and Chunked Causal Convolution
di: Wang, Fangyuan, et al.
Pubblicazione: (2022)
di: Wang, Fangyuan, et al.
Pubblicazione: (2022)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
di: Nakagome, Yu, et al.
Pubblicazione: (2025)
di: Nakagome, Yu, et al.
Pubblicazione: (2025)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
Efficient Scaling for LLM-based ASR
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
SOT Triggered Neural Clustering for Speaker Attributed ASR
di: Zheng, Xianrui, et al.
Pubblicazione: (2024)
di: Zheng, Xianrui, et al.
Pubblicazione: (2024)
DNCASR: End-to-End Training for Speaker-Attributed ASR
di: Zheng, Xianrui, et al.
Pubblicazione: (2025)
di: Zheng, Xianrui, et al.
Pubblicazione: (2025)
An investigation of modularity for noise robustness in conformer-based ASR
di: de Gibson, Louise Coppieters, et al.
Pubblicazione: (2024)
di: de Gibson, Louise Coppieters, et al.
Pubblicazione: (2024)
Expressive Timing in Hindustani Vocal Music
di: Bhake, Yash, et al.
Pubblicazione: (2025)
di: Bhake, Yash, et al.
Pubblicazione: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2024)
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2024)
NLE: Non-autoregressive LLM-based ASR by Transcript Editing
di: Dekel, Avihu, et al.
Pubblicazione: (2026)
di: Dekel, Avihu, et al.
Pubblicazione: (2026)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
Towards a Single ASR Model That Generalizes to Disordered Speech
di: Tobin, Jimmy, et al.
Pubblicazione: (2024)
di: Tobin, Jimmy, et al.
Pubblicazione: (2024)
The USTC-NERCSLIP Systems for The ICMC-ASR Challenge
di: Wu, Minghui, et al.
Pubblicazione: (2024)
di: Wu, Minghui, et al.
Pubblicazione: (2024)
AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
di: Gündüz, Ahmet, et al.
Pubblicazione: (2024)
di: Gündüz, Ahmet, et al.
Pubblicazione: (2024)
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
di: Wu, Ke, et al.
Pubblicazione: (2026)
di: Wu, Ke, et al.
Pubblicazione: (2026)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
di: Tsai, Kai-Chen, et al.
Pubblicazione: (2026)
di: Tsai, Kai-Chen, et al.
Pubblicazione: (2026)
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
di: Fujita, Yuya, et al.
Pubblicazione: (2024)
di: Fujita, Yuya, et al.
Pubblicazione: (2024)
Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment
di: Rousso, Rotem, et al.
Pubblicazione: (2024)
di: Rousso, Rotem, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Anatomy of Industrial Scale Multilingual ASR
di: Ramirez, Francis McCann, et al.
Pubblicazione: (2024) -
Universal-2-TF: Robust All-Neural Text Formatting for ASR
di: Khare, Yash, et al.
Pubblicazione: (2025) -
An open-source voice type classifier for child-centered daylong recordings
di: Lavechin, Marvin, et al.
Pubblicazione: (2020) -
Promptformer: Prompted Conformer Transducer for ASR
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024) -
Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation
di: Cai, Danwei, et al.
Pubblicazione: (2023)