Unimodal Aggregation for CTC-based Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Ying, Li, Xiaofei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mamba for Streaming ASR Combined with Unimodal Aggregation
von: Fang, Ying, et al.
Veröffentlicht: (2024)
von: Fang, Ying, et al.
Veröffentlicht: (2024)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
von: Abu, Turi, et al.
Veröffentlicht: (2025)
von: Abu, Turi, et al.
Veröffentlicht: (2025)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
von: Zheng, Yuang, et al.
Veröffentlicht: (2026)
von: Zheng, Yuang, et al.
Veröffentlicht: (2026)
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
von: Zhang, Tian-Hao, et al.
Veröffentlicht: (2023)
von: Zhang, Tian-Hao, et al.
Veröffentlicht: (2023)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition for Hindi
von: Saha, Anish, et al.
Veröffentlicht: (2024)
von: Saha, Anish, et al.
Veröffentlicht: (2024)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
von: Zhang, Enshi, et al.
Veröffentlicht: (2024)
von: Zhang, Enshi, et al.
Veröffentlicht: (2024)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
von: Altinok, Duygu
Veröffentlicht: (2025)
von: Altinok, Duygu
Veröffentlicht: (2025)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
von: Nakagome, Yu, et al.
Veröffentlicht: (2025)
von: Nakagome, Yu, et al.
Veröffentlicht: (2025)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
von: Hu, Jiliang, et al.
Veröffentlicht: (2024)
von: Hu, Jiliang, et al.
Veröffentlicht: (2024)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Dynamic Data Pruning for Automatic Speech Recognition
von: Xiao, Qiao, et al.
Veröffentlicht: (2024)
von: Xiao, Qiao, et al.
Veröffentlicht: (2024)
On the Contribution of Lexical Features to Speech Emotion Recognition
von: Combei, David
Veröffentlicht: (2025)
von: Combei, David
Veröffentlicht: (2025)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Continual Adaptation for Pacific Indigenous Speech Recognition
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
Dialectal Coverage And Generalization in Arabic Speech Recognition
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2024)
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2024)
Speech Recognition Transformers: Topological-lingualism Perspective
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
DiariST: Streaming Speech Translation with Speaker Diarization
von: Yang, Mu, et al.
Veröffentlicht: (2023)
von: Yang, Mu, et al.
Veröffentlicht: (2023)
Enhancing CTC-Based Visual Speech Recognition
von: Laux, Hendrik, et al.
Veröffentlicht: (2024)
von: Laux, Hendrik, et al.
Veröffentlicht: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition
von: Xia, Yinfeng, et al.
Veröffentlicht: (2025)
von: Xia, Yinfeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mamba for Streaming ASR Combined with Unimodal Aggregation
von: Fang, Ying, et al.
Veröffentlicht: (2024) -
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025) -
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024) -
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
von: Ma, Rao, et al.
Veröffentlicht: (2025) -
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)