A two-stage transliteration approach to improve performance of a multilingual ASR
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Kumar, Rohit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR
von: Mei, Yuxiang, et al.
Veröffentlicht: (2026)
von: Mei, Yuxiang, et al.
Veröffentlicht: (2026)
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
von: Kopparapu, Sunil Kumar, et al.
Veröffentlicht: (2024)
von: Kopparapu, Sunil Kumar, et al.
Veröffentlicht: (2024)
Quantizing Whisper-small: How design choices affect ASR performance
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2024)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2024)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
von: Kumar, Sujeet, et al.
Veröffentlicht: (2025)
von: Kumar, Sujeet, et al.
Veröffentlicht: (2025)
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
von: Sivasankaran, Sunit, et al.
Veröffentlicht: (2024)
von: Sivasankaran, Sunit, et al.
Veröffentlicht: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
ZIPA: A family of efficient models for multilingual phone recognition
von: Zhu, Jian, et al.
Veröffentlicht: (2025)
von: Zhu, Jian, et al.
Veröffentlicht: (2025)
AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2024)
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2024)
Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices
von: Velikovich, Leonid, et al.
Veröffentlicht: (2024)
von: Velikovich, Leonid, et al.
Veröffentlicht: (2024)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
Romanization Encoding For Multilingual ASR
von: Ding, Wen, et al.
Veröffentlicht: (2024)
von: Ding, Wen, et al.
Veröffentlicht: (2024)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
Promptformer: Prompted Conformer Transducer for ASR
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
Revisiting Acoustic Features for Robust ASR
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Qwen3-ASR Technical Report
von: Shi, Xian, et al.
Veröffentlicht: (2026)
von: Shi, Xian, et al.
Veröffentlicht: (2026)
ASR Benchmarking: Need for a More Representative Conversational Dataset
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
von: Kumar, Subham, et al.
Veröffentlicht: (2025)
von: Kumar, Subham, et al.
Veröffentlicht: (2025)
Exploring SSL Discrete Tokens for Multilingual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Configurable Multilingual ASR with Speech Summary Representations
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
von: Zhu, Harrison, et al.
Veröffentlicht: (2024)
ManWav: The First Manchu ASR Model
von: Seo, Jean, et al.
Veröffentlicht: (2024)
von: Seo, Jean, et al.
Veröffentlicht: (2024)
Mamba for Streaming ASR Combined with Unimodal Aggregation
von: Fang, Ying, et al.
Veröffentlicht: (2024)
von: Fang, Ying, et al.
Veröffentlicht: (2024)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
HypR: A comprehensive study for ASR hypothesis revising with a reference corpus
von: Wang, Yi-Wei, et al.
Veröffentlicht: (2023)
von: Wang, Yi-Wei, et al.
Veröffentlicht: (2023)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
von: Wang, He, et al.
Veröffentlicht: (2025)
von: Wang, He, et al.
Veröffentlicht: (2025)
Performant ASR Models for Medical Entities in Accented Speech
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
Reverb: Open-Source ASR and Diarization from Rev
von: Bhandari, Nishchal, et al.
Veröffentlicht: (2024)
von: Bhandari, Nishchal, et al.
Veröffentlicht: (2024)
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
WER We Stand: Benchmarking Urdu ASR Models
von: Arif, Samee, et al.
Veröffentlicht: (2024)
von: Arif, Samee, et al.
Veröffentlicht: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Advocating Character Error Rate for Multilingual ASR Evaluation
von: K, Thennal D, et al.
Veröffentlicht: (2024)
von: K, Thennal D, et al.
Veröffentlicht: (2024)
Scalable Offline ASR for Command-Style Dictation in Courtrooms
von: Nethil, Kumarmanas, et al.
Veröffentlicht: (2025)
von: Nethil, Kumarmanas, et al.
Veröffentlicht: (2025)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
Causal Structure Discovery for Error Diagnostics of Children's ASR
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR
von: Mei, Yuxiang, et al.
Veröffentlicht: (2026) -
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
von: Kopparapu, Sunil Kumar, et al.
Veröffentlicht: (2024) -
Quantizing Whisper-small: How design choices affect ASR performance
von: Söhler, Arthur, et al.
Veröffentlicht: (2025) -
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2024) -
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)