Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Polok, Alexander, Kesiraju, Santosh, Beneš, Karel, Burget, Lukáš, Černocký, Jan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets
von: Pešán, Jan, et al.
Veröffentlicht: (2024)
von: Pešán, Jan, et al.
Veröffentlicht: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
BUT System for the MLC-SLM Challenge
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
Unsupervised Speech Enhancement using Data-defined Priors
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
Modeling Overlapped Speech with Shuffles
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech Separation
von: Shin, Ui-Hyeop, et al.
Veröffentlicht: (2024)
von: Shin, Ui-Hyeop, et al.
Veröffentlicht: (2024)
BUT System Description for CHiME-9 MCoRec Challenge
von: Klement, Dominik, et al.
Veröffentlicht: (2026)
von: Klement, Dominik, et al.
Veröffentlicht: (2026)
State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
von: Barahona, Sara, et al.
Veröffentlicht: (2024)
von: Barahona, Sara, et al.
Veröffentlicht: (2024)
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
von: Tawara, Naohiro, et al.
Veröffentlicht: (2026)
von: Tawara, Naohiro, et al.
Veröffentlicht: (2026)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
Training and Inference Efficiency of Encoder-Decoder Speech Models
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
von: Cohen, Eyal, et al.
Veröffentlicht: (2025)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
von: Wei, Linye, et al.
Veröffentlicht: (2025)
von: Wei, Linye, et al.
Veröffentlicht: (2025)
Asymmetric Encoder-Decoder Based on Time-Frequency Correlation for Speech Separation
von: Shin, Ui-Hyeop, et al.
Veröffentlicht: (2026)
von: Shin, Ui-Hyeop, et al.
Veröffentlicht: (2026)
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
von: Xu, Kai-Tuo, et al.
Veröffentlicht: (2025)
von: Xu, Kai-Tuo, et al.
Veröffentlicht: (2025)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
Using Songs to Improve Kazakh Automatic Speech Recognition
von: Yeshpanov, Rustem
Veröffentlicht: (2026)
von: Yeshpanov, Rustem
Veröffentlicht: (2026)
Speech-Aware Neural Diarization with Encoder-Decoder Attractor Guided by Attention Constraints
von: Lee, PeiYing, et al.
Veröffentlicht: (2024)
von: Lee, PeiYing, et al.
Veröffentlicht: (2024)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
von: Wan, Genshun, et al.
Veröffentlicht: (2026)
von: Wan, Genshun, et al.
Veröffentlicht: (2026)
Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
von: Jinnai, Yuu
Veröffentlicht: (2025)
von: Jinnai, Yuu
Veröffentlicht: (2025)
Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
von: Saon, George, et al.
Veröffentlicht: (2026)
von: Saon, George, et al.
Veröffentlicht: (2026)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
Robustness assessment of large audio language models in multiple-choice evaluation
von: López, Fernando, et al.
Veröffentlicht: (2025)
von: López, Fernando, et al.
Veröffentlicht: (2025)
Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement
von: Haider, Daniel, et al.
Veröffentlicht: (2024)
von: Haider, Daniel, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2025) -
Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech Recognition Datasets
von: Pešán, Jan, et al.
Veröffentlicht: (2024) -
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025) -
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024) -
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
von: Polok, Alexander, et al.
Veröffentlicht: (2026)