Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
Fuente:
arXiv
Salvato in:
| Autori principali: | Kang, Wei, Yang, Xiaoyu, Yao, Zengwei, Kuang, Fangjun, Yang, Yifan, Guo, Liyong, Lin, Long, Povey, Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PromptASR for contextualized ASR with controllable style
di: Yang, Xiaoyu, et al.
Pubblicazione: (2023)
di: Yang, Xiaoyu, et al.
Pubblicazione: (2023)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
di: Jin, Zengrui, et al.
Pubblicazione: (2024)
di: Jin, Zengrui, et al.
Pubblicazione: (2024)
Zipformer: A faster and better encoder for automatic speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2023)
di: Yao, Zengwei, et al.
Pubblicazione: (2023)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
di: Zhu, Han, et al.
Pubblicazione: (2025)
di: Zhu, Han, et al.
Pubblicazione: (2025)
CR-CTC: Consistency regularization on CTC for improved speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
di: Yang, Yifan, et al.
Pubblicazione: (2024)
di: Yang, Yifan, et al.
Pubblicazione: (2024)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
di: Zhuo, Jianheng, et al.
Pubblicazione: (2025)
di: Zhuo, Jianheng, et al.
Pubblicazione: (2025)
Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
di: Yao, Zengwei, et al.
Pubblicazione: (2025)
di: Yao, Zengwei, et al.
Pubblicazione: (2025)
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
di: You, Jian, et al.
Pubblicazione: (2024)
di: You, Jian, et al.
Pubblicazione: (2024)
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
di: Zhu, Han, et al.
Pubblicazione: (2026)
di: Zhu, Han, et al.
Pubblicazione: (2026)
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
di: Chen, Guoguo, et al.
Pubblicazione: (2021)
di: Chen, Guoguo, et al.
Pubblicazione: (2021)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
di: Li, Shaojun, et al.
Pubblicazione: (2024)
di: Li, Shaojun, et al.
Pubblicazione: (2024)
Index-ASR Technical Report
di: Song, Zheshu, et al.
Pubblicazione: (2025)
di: Song, Zheshu, et al.
Pubblicazione: (2025)
DQLoRA: A Lightweight Domain-Aware Denoising ASR via Adapter-guided Distillation
di: Yang, Yiru
Pubblicazione: (2025)
di: Yang, Yiru
Pubblicazione: (2025)
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
di: Xu, Kai-Tuo, et al.
Pubblicazione: (2025)
di: Xu, Kai-Tuo, et al.
Pubblicazione: (2025)
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
LUPET: Incorporating Hierarchical Information Path into Multilingual ASR
di: Liu, Wei, et al.
Pubblicazione: (2024)
di: Liu, Wei, et al.
Pubblicazione: (2024)
Towards Decoupling Frontend Enhancement and Backend Recognition in Monaural Robust ASR
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
di: Pulikodan, Sujith, et al.
Pubblicazione: (2025)
di: Pulikodan, Sujith, et al.
Pubblicazione: (2025)
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
di: Yang, Yufeng, et al.
Pubblicazione: (2025)
di: Yang, Yufeng, et al.
Pubblicazione: (2025)
Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification
di: Gao, Yifan, et al.
Pubblicazione: (2024)
di: Gao, Yifan, et al.
Pubblicazione: (2024)
HypR: A comprehensive study for ASR hypothesis revising with a reference corpus
di: Wang, Yi-Wei, et al.
Pubblicazione: (2023)
di: Wang, Yi-Wei, et al.
Pubblicazione: (2023)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
di: Yang, Yufeng, et al.
Pubblicazione: (2025)
di: Yang, Yufeng, et al.
Pubblicazione: (2025)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
di: Xu, Kaituo, et al.
Pubblicazione: (2026)
di: Xu, Kaituo, et al.
Pubblicazione: (2026)
Exploring SSL Discrete Tokens for Multilingual ASR
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
di: Cui, Mingyu, et al.
Pubblicazione: (2024)
Target Speaker ASR with Whisper
di: Polok, Alexander, et al.
Pubblicazione: (2024)
di: Polok, Alexander, et al.
Pubblicazione: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
di: Guo, Pengcheng, et al.
Pubblicazione: (2024)
di: Guo, Pengcheng, et al.
Pubblicazione: (2024)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
Efficient Scaling for LLM-based ASR
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
On Speaker Attribution with SURT
di: Raj, Desh, et al.
Pubblicazione: (2024)
di: Raj, Desh, et al.
Pubblicazione: (2024)
The USTC-NERCSLIP Systems for The ICMC-ASR Challenge
di: Wu, Minghui, et al.
Pubblicazione: (2024)
di: Wu, Minghui, et al.
Pubblicazione: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
di: Bai, Ye, et al.
Pubblicazione: (2024)
di: Bai, Ye, et al.
Pubblicazione: (2024)
Qwen3-ASR Technical Report
di: Shi, Xian, et al.
Pubblicazione: (2026)
di: Shi, Xian, et al.
Pubblicazione: (2026)
persoDA: Personalized Data Augmentation for Personalized ASR
di: Parada, Pablo Peso, et al.
Pubblicazione: (2025)
di: Parada, Pablo Peso, et al.
Pubblicazione: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
Comparative Analysis of ASR Methods for Speech Deepfake Detection
di: Salvi, Davide, et al.
Pubblicazione: (2024)
di: Salvi, Davide, et al.
Pubblicazione: (2024)
Consistency Based Unsupervised Self-training For ASR Personalisation
di: Zhang, Jisi, et al.
Pubblicazione: (2024)
di: Zhang, Jisi, et al.
Pubblicazione: (2024)
THAI Speech Emotion Recognition (THAI-SER) corpus
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
di: Chen, Szu-Jui, et al.
Pubblicazione: (2026)
di: Chen, Szu-Jui, et al.
Pubblicazione: (2026)
Documenti analoghi
-
PromptASR for contextualized ASR with controllable style
di: Yang, Xiaoyu, et al.
Pubblicazione: (2023) -
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
di: Jin, Zengrui, et al.
Pubblicazione: (2024) -
Zipformer: A faster and better encoder for automatic speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2023) -
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
di: Zhu, Han, et al.
Pubblicazione: (2025) -
CR-CTC: Consistency regularization on CTC for improved speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2024)