Salvato in:
| Autori principali: | Xu, Jingjing, Zhou, Wei, Yang, Zijian, Beck, Eugen, Schlueter, Ralf |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2407.18930 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dynamic Data Pruning for Automatic Speech Recognition
di: Xiao, Qiao, et al.
Pubblicazione: (2024)
di: Xiao, Qiao, et al.
Pubblicazione: (2024)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025)
Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations
di: He, Linyang, et al.
Pubblicazione: (2025)
di: He, Linyang, et al.
Pubblicazione: (2025)
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
di: Choi, Anna Seo Gyeong, et al.
Pubblicazione: (2025)
di: Choi, Anna Seo Gyeong, et al.
Pubblicazione: (2025)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
di: Lin, Zhennan, et al.
Pubblicazione: (2025)
di: Lin, Zhennan, et al.
Pubblicazione: (2025)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
Speech-Based Depression Prediction Using Encoder-Weight-Only Transfer Learning and a Large Corpus
di: Harati, Amir, et al.
Pubblicazione: (2024)
di: Harati, Amir, et al.
Pubblicazione: (2024)
An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement
di: Yang, Tzu-Ting, et al.
Pubblicazione: (2024)
di: Yang, Tzu-Ting, et al.
Pubblicazione: (2024)
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
di: Li, Bohan, et al.
Pubblicazione: (2025)
di: Li, Bohan, et al.
Pubblicazione: (2025)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2024)
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2024)
Training and Inference Efficiency of Encoder-Decoder Speech Models
di: Żelasko, Piotr, et al.
Pubblicazione: (2025)
di: Żelasko, Piotr, et al.
Pubblicazione: (2025)
Task-Agnostic Structured Pruning of Speech Representation Models
di: Wang, Haoyu, et al.
Pubblicazione: (2023)
di: Wang, Haoyu, et al.
Pubblicazione: (2023)
SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
di: Du, Jiayu, et al.
Pubblicazione: (2024)
di: Du, Jiayu, et al.
Pubblicazione: (2024)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
di: Rossenbach, Nick, et al.
Pubblicazione: (2024)
di: Rossenbach, Nick, et al.
Pubblicazione: (2024)
Data Augmentation for End-to-end Code-switching Speech Recognition
di: Du, Chenpeng, et al.
Pubblicazione: (2020)
di: Du, Chenpeng, et al.
Pubblicazione: (2020)
Context-Driven Dynamic Pruning for Large Speech Foundation Models
di: Someki, Masao, et al.
Pubblicazione: (2025)
di: Someki, Masao, et al.
Pubblicazione: (2025)
Streaming Speech-to-Confusion Network Speech Recognition
di: Filimonov, Denis, et al.
Pubblicazione: (2023)
di: Filimonov, Denis, et al.
Pubblicazione: (2023)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
di: Sudo, Yui, et al.
Pubblicazione: (2024)
di: Sudo, Yui, et al.
Pubblicazione: (2024)
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
di: Shao, Qijie, et al.
Pubblicazione: (2025)
di: Shao, Qijie, et al.
Pubblicazione: (2025)
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
di: Sedláček, Šimon, et al.
Pubblicazione: (2025)
di: Sedláček, Šimon, et al.
Pubblicazione: (2025)
Advancing African-Accented Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models
di: Dossou, Bonaventure F. P.
Pubblicazione: (2023)
di: Dossou, Bonaventure F. P.
Pubblicazione: (2023)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
di: Xie, Yuan, et al.
Pubblicazione: (2026)
di: Xie, Yuan, et al.
Pubblicazione: (2026)
Convexity-based Pruning of Speech Representation Models
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
di: Polok, Alexander, et al.
Pubblicazione: (2025)
di: Polok, Alexander, et al.
Pubblicazione: (2025)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
di: Yang, Zijian, et al.
Pubblicazione: (2023)
di: Yang, Zijian, et al.
Pubblicazione: (2023)
Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
di: Matsuura, Kohei, et al.
Pubblicazione: (2024)
di: Matsuura, Kohei, et al.
Pubblicazione: (2024)
Automatic Speech Recognition for Biomedical Data in Bengali Language
di: Kabir, Shariar, et al.
Pubblicazione: (2024)
di: Kabir, Shariar, et al.
Pubblicazione: (2024)
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
di: Wang, Tianduo, et al.
Pubblicazione: (2025)
di: Wang, Tianduo, et al.
Pubblicazione: (2025)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
di: Hilmes, Benedikt, et al.
Pubblicazione: (2024)
di: Hilmes, Benedikt, et al.
Pubblicazione: (2024)
Chain of Correction for Full-text Speech Recognition with Large Language Models
di: Tang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Tang, Zhiyuan, et al.
Pubblicazione: (2025)
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
di: Yang, Runyan, et al.
Pubblicazione: (2025)
di: Yang, Runyan, et al.
Pubblicazione: (2025)
AdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge Integration
di: Lee, Chia-Yu, et al.
Pubblicazione: (2026)
di: Lee, Chia-Yu, et al.
Pubblicazione: (2026)
Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
di: Sameti, Mohammad Hossein, et al.
Pubblicazione: (2025)
di: Sameti, Mohammad Hossein, et al.
Pubblicazione: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
Frontend Token Enhancement for Token-Based Speech Recognition
di: Ashihara, Takanori, et al.
Pubblicazione: (2026)
di: Ashihara, Takanori, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Dynamic Data Pruning for Automatic Speech Recognition
di: Xiao, Qiao, et al.
Pubblicazione: (2024) -
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025) -
Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations
di: He, Linyang, et al.
Pubblicazione: (2025) -
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
di: Choi, Anna Seo Gyeong, et al.
Pubblicazione: (2025) -
Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
di: Lin, Zhennan, et al.
Pubblicazione: (2025)