A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | You, Jian, Li, Xiangfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
von: Sivasankaran, Sunit, et al.
Veröffentlicht: (2024)
von: Sivasankaran, Sunit, et al.
Veröffentlicht: (2024)
Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
von: Bai, Junwen, et al.
Veröffentlicht: (2024)
von: Bai, Junwen, et al.
Veröffentlicht: (2024)
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
von: Kang, Wei, et al.
Veröffentlicht: (2023)
von: Kang, Wei, et al.
Veröffentlicht: (2023)
Anatomy of Industrial Scale Multilingual ASR
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
Beyond Transcription: Mechanistic Interpretability in ASR
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
SALSA: Speedy ASR-LLM Synchronous Aggregation
von: Mittal, Ashish, et al.
Veröffentlicht: (2024)
von: Mittal, Ashish, et al.
Veröffentlicht: (2024)
Revisiting ASR Error Correction with Specialized Models
von: Gu, Zijin, et al.
Veröffentlicht: (2024)
von: Gu, Zijin, et al.
Veröffentlicht: (2024)
Federated Learning of Large ASR Models in the Real World
von: Xiao, Yonghui, et al.
Veröffentlicht: (2024)
von: Xiao, Yonghui, et al.
Veröffentlicht: (2024)
Unsupervised ASR via Cross-Lingual Pseudo-Labeling
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2023)
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2023)
Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
von: Galimzianov, Dmitrii, et al.
Veröffentlicht: (2024)
von: Galimzianov, Dmitrii, et al.
Veröffentlicht: (2024)
Transformer-based Model for ASR N-Best Rescoring and Rewriting
von: Kang, Iwen E., et al.
Veröffentlicht: (2024)
von: Kang, Iwen E., et al.
Veröffentlicht: (2024)
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
von: Lei, Zhihong, et al.
Veröffentlicht: (2024)
von: Lei, Zhihong, et al.
Veröffentlicht: (2024)
Cross-utterance ASR Rescoring with Graph-based Label Propagation
von: Tankasala, Srinath, et al.
Veröffentlicht: (2023)
von: Tankasala, Srinath, et al.
Veröffentlicht: (2023)
Unified Learnable 2D Convolutional Feature Extraction for ASR
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping
von: Zhang, Kevin, et al.
Veröffentlicht: (2024)
von: Zhang, Kevin, et al.
Veröffentlicht: (2024)
Text-only adaptation in LLM-based ASR through text denoising
von: Carofilis, Andrés, et al.
Veröffentlicht: (2026)
von: Carofilis, Andrés, et al.
Veröffentlicht: (2026)
Dysarthria Normalization via Local Lie Group Transformations for Robust ASR
von: Osipov, Mikhail
Veröffentlicht: (2025)
von: Osipov, Mikhail
Veröffentlicht: (2025)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
von: Dissen, Yehoshua, et al.
Veröffentlicht: (2024)
von: Dissen, Yehoshua, et al.
Veröffentlicht: (2024)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
Benchmarking Akan ASR Models Across Domain-Specific Datasets: A Comparative Evaluation of Performance, Scalability, and Adaptability
von: Mensah, Mark Atta, et al.
Veröffentlicht: (2025)
von: Mensah, Mark Atta, et al.
Veröffentlicht: (2025)
A low latency attention module for streaming self-supervised speech representation learning
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese
von: Arisaputra, Panji, et al.
Veröffentlicht: (2024)
von: Arisaputra, Panji, et al.
Veröffentlicht: (2024)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
von: Bataev, Vladimir, et al.
Veröffentlicht: (2023)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2023)
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
von: Orhon, Atila, et al.
Veröffentlicht: (2025)
von: Orhon, Atila, et al.
Veröffentlicht: (2025)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
von: Amooie, Reihaneh, et al.
Veröffentlicht: (2025)
von: Amooie, Reihaneh, et al.
Veröffentlicht: (2025)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
An ASR-Based Tutor for Learning to Read: How to Optimize Feedback to First Graders
von: Bai, Yu, et al.
Veröffentlicht: (2023)
von: Bai, Yu, et al.
Veröffentlicht: (2023)
Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
von: Cuervo, Santiago, et al.
Veröffentlicht: (2021)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2021)
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
Selective Attention Merging for low resource tasks: A case study of Child ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
ZIPA: A family of efficient models for multilingual phone recognition
von: Zhu, Jian, et al.
Veröffentlicht: (2025)
von: Zhu, Jian, et al.
Veröffentlicht: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
Mamba for Streaming ASR Combined with Unimodal Aggregation
von: Fang, Ying, et al.
Veröffentlicht: (2024)
von: Fang, Ying, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
von: Sivasankaran, Sunit, et al.
Veröffentlicht: (2024) -
Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
von: Bai, Junwen, et al.
Veröffentlicht: (2024) -
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
von: Kang, Wei, et al.
Veröffentlicht: (2023) -
Anatomy of Industrial Scale Multilingual ASR
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024) -
Beyond Transcription: Mechanistic Interpretability in ASR
von: Glazer, Neta, et al.
Veröffentlicht: (2025)