ASR Benchmarking: Need for a More Representative Conversational Dataset
Fuente:
arXiv
Guardado en:
| Autores principales: | Maheshwari, Gaurav, Ivanov, Dmitry, Johannet, Théo, Haddad, Kevin El |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
por: Wei, Victor Junqiu, et al.
Publicado: (2024)
por: Wei, Victor Junqiu, et al.
Publicado: (2024)
WER We Stand: Benchmarking Urdu ASR Models
por: Arif, Samee, et al.
Publicado: (2024)
por: Arif, Samee, et al.
Publicado: (2024)
LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization
por: Gedeon, Máté, et al.
Publicado: (2025)
por: Gedeon, Máté, et al.
Publicado: (2025)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
por: Chen, Qian, et al.
Publicado: (2023)
por: Chen, Qian, et al.
Publicado: (2023)
How Much Context Does My Attention-Based ASR System Need?
por: Flynn, Robert, et al.
Publicado: (2023)
por: Flynn, Robert, et al.
Publicado: (2023)
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
por: Kumar, Sujeet, et al.
Publicado: (2025)
por: Kumar, Sujeet, et al.
Publicado: (2025)
Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
por: Geng, Xuelong, et al.
Publicado: (2024)
por: Geng, Xuelong, et al.
Publicado: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
por: Fan, Ruchao, et al.
Publicado: (2024)
por: Fan, Ruchao, et al.
Publicado: (2024)
ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems
por: Rai, Anand, et al.
Publicado: (2025)
por: Rai, Anand, et al.
Publicado: (2025)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
por: Wang, He, et al.
Publicado: (2025)
por: Wang, He, et al.
Publicado: (2025)
PromptASR for contextualized ASR with controllable style
por: Yang, Xiaoyu, et al.
Publicado: (2023)
por: Yang, Xiaoyu, et al.
Publicado: (2023)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
por: Li, Chin-Jou, et al.
Publicado: (2025)
por: Li, Chin-Jou, et al.
Publicado: (2025)
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
por: Kamble, Anand, et al.
Publicado: (2023)
por: Kamble, Anand, et al.
Publicado: (2023)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
por: Gündüz, Ahmet, et al.
Publicado: (2024)
por: Gündüz, Ahmet, et al.
Publicado: (2024)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
por: Ko, Yuka, et al.
Publicado: (2024)
por: Ko, Yuka, et al.
Publicado: (2024)
Romanization Encoding For Multilingual ASR
por: Ding, Wen, et al.
Publicado: (2024)
por: Ding, Wen, et al.
Publicado: (2024)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
por: Xie, Yuan, et al.
Publicado: (2026)
por: Xie, Yuan, et al.
Publicado: (2026)
Promptformer: Prompted Conformer Transducer for ASR
por: Duarte-Torres, Sergio, et al.
Publicado: (2024)
por: Duarte-Torres, Sergio, et al.
Publicado: (2024)
Revisiting Acoustic Features for Robust ASR
por: Shah, Muhammad A., et al.
Publicado: (2024)
por: Shah, Muhammad A., et al.
Publicado: (2024)
Qwen3-ASR Technical Report
por: Shi, Xian, et al.
Publicado: (2026)
por: Shi, Xian, et al.
Publicado: (2026)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
por: Jin, Zengrui, et al.
Publicado: (2024)
por: Jin, Zengrui, et al.
Publicado: (2024)
Exploring SSL Discrete Tokens for Multilingual ASR
por: Cui, Mingyu, et al.
Publicado: (2024)
por: Cui, Mingyu, et al.
Publicado: (2024)
Configurable Multilingual ASR with Speech Summary Representations
por: Zhu, Harrison, et al.
Publicado: (2024)
por: Zhu, Harrison, et al.
Publicado: (2024)
ManWav: The First Manchu ASR Model
por: Seo, Jean, et al.
Publicado: (2024)
por: Seo, Jean, et al.
Publicado: (2024)
Mamba for Streaming ASR Combined with Unimodal Aggregation
por: Fang, Ying, et al.
Publicado: (2024)
por: Fang, Ying, et al.
Publicado: (2024)
Semi-Autoregressive Streaming ASR With Label Context
por: Arora, Siddhant, et al.
Publicado: (2023)
por: Arora, Siddhant, et al.
Publicado: (2023)
Performant ASR Models for Medical Entities in Accented Speech
por: Afonja, Tejumade, et al.
Publicado: (2024)
por: Afonja, Tejumade, et al.
Publicado: (2024)
Reverb: Open-Source ASR and Diarization from Rev
por: Bhandari, Nishchal, et al.
Publicado: (2024)
por: Bhandari, Nishchal, et al.
Publicado: (2024)
ASR Error Correction using Large Language Models
por: Ma, Rao, et al.
Publicado: (2024)
por: Ma, Rao, et al.
Publicado: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
Advocating Character Error Rate for Multilingual ASR Evaluation
por: K, Thennal D, et al.
Publicado: (2024)
por: K, Thennal D, et al.
Publicado: (2024)
Scalable Offline ASR for Command-Style Dictation in Courtrooms
por: Nethil, Kumarmanas, et al.
Publicado: (2025)
por: Nethil, Kumarmanas, et al.
Publicado: (2025)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
por: Shakeel, Muhammad, et al.
Publicado: (2025)
por: Shakeel, Muhammad, et al.
Publicado: (2025)
Causal Structure Discovery for Error Diagnostics of Children's ASR
por: Singh, Vishwanath Pratap, et al.
Publicado: (2025)
por: Singh, Vishwanath Pratap, et al.
Publicado: (2025)
Extending Whisper with prompt tuning to target-speaker ASR
por: Ma, Hao, et al.
Publicado: (2023)
por: Ma, Hao, et al.
Publicado: (2023)
A two-stage transliteration approach to improve performance of a multilingual ASR
por: Kumar, Rohit
Publicado: (2024)
por: Kumar, Rohit
Publicado: (2024)
Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices
por: Velikovich, Leonid, et al.
Publicado: (2024)
por: Velikovich, Leonid, et al.
Publicado: (2024)
ProGRes: Prompted Generative Rescoring on ASR n-Best
por: Tur, Ada Defne, et al.
Publicado: (2024)
por: Tur, Ada Defne, et al.
Publicado: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
por: Moriya, Takafumi, et al.
Publicado: (2024)
por: Moriya, Takafumi, et al.
Publicado: (2024)
Ejemplares similares
-
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
por: Wei, Victor Junqiu, et al.
Publicado: (2024) -
WER We Stand: Benchmarking Urdu ASR Models
por: Arif, Samee, et al.
Publicado: (2024) -
LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization
por: Gedeon, Máté, et al.
Publicado: (2025) -
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
por: Chen, Qian, et al.
Publicado: (2023) -
How Much Context Does My Attention-Based ASR System Need?
por: Flynn, Robert, et al.
Publicado: (2023)