Combining X-Vectors and Bayesian Batch Active Learning: Two-Stage Active Learning Pipeline for Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kundacina, Ognjen, Vincan, Vladimir, Miskovic, Dragisa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
Active Learning with Task Adaptation Pre-training for Speech Emotion Recognition
von: Li, Dongyuan, et al.
Veröffentlicht: (2024)
von: Li, Dongyuan, et al.
Veröffentlicht: (2024)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
von: Bataev, Vladimir
Veröffentlicht: (2025)
von: Bataev, Vladimir
Veröffentlicht: (2025)
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
von: Yaish, Ofir, et al.
Veröffentlicht: (2025)
von: Yaish, Ofir, et al.
Veröffentlicht: (2025)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
von: Storey, Edward, et al.
Veröffentlicht: (2025)
von: Storey, Edward, et al.
Veröffentlicht: (2025)
Pretraining Large Brain Language Model for Active BCI: Silent Speech
von: Zhou, Jinzhao, et al.
Veröffentlicht: (2025)
von: Zhou, Jinzhao, et al.
Veröffentlicht: (2025)
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
von: Hono, Yukiya, et al.
Veröffentlicht: (2023)
von: Hono, Yukiya, et al.
Veröffentlicht: (2023)
TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
Semantically Corrected Amharic Automatic Speech Recognition
von: Adnew, Samuael, et al.
Veröffentlicht: (2024)
von: Adnew, Samuael, et al.
Veröffentlicht: (2024)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
von: Chen, William, et al.
Veröffentlicht: (2025)
von: Chen, William, et al.
Veröffentlicht: (2025)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
von: Min, Do June, et al.
Veröffentlicht: (2024)
von: Min, Do June, et al.
Veröffentlicht: (2024)
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
von: Zhong, Tao, et al.
Veröffentlicht: (2025)
von: Zhong, Tao, et al.
Veröffentlicht: (2025)
Handling Numeric Expressions in Automatic Speech Recognition
von: Huber, Christian, et al.
Veröffentlicht: (2024)
von: Huber, Christian, et al.
Veröffentlicht: (2024)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
von: Zhang, Shucong, et al.
Veröffentlicht: (2025)
von: Zhang, Shucong, et al.
Veröffentlicht: (2025)
Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition
von: Vu, Tai
Veröffentlicht: (2025)
von: Vu, Tai
Veröffentlicht: (2025)
Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2023)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2023)
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
von: Billa, Jayadev
Veröffentlicht: (2026)
von: Billa, Jayadev
Veröffentlicht: (2026)
Reading Miscue Detection in Primary School through Automatic Speech Recognition
von: Gao, Lingyun, et al.
Veröffentlicht: (2024)
von: Gao, Lingyun, et al.
Veröffentlicht: (2024)
Towards Robust Speech Recognition for Jamaican Patois Music Transcription
von: Madden, Jordan, et al.
Veröffentlicht: (2025)
von: Madden, Jordan, et al.
Veröffentlicht: (2025)
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
von: Vu, Tai
Veröffentlicht: (2025)
von: Vu, Tai
Veröffentlicht: (2025)
Deep Active Speech Cancellation with Mamba-Masking Network
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2024)
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2024)
Linear-Complexity Self-Supervised Learning for Speech Processing
von: Zhang, Shucong, et al.
Veröffentlicht: (2024)
von: Zhang, Shucong, et al.
Veröffentlicht: (2024)
Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge
von: Li, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Li, Xiaoxiao, et al.
Veröffentlicht: (2025)
SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
von: Chang, Kai-Wei, et al.
Veröffentlicht: (2024)
von: Chang, Kai-Wei, et al.
Veröffentlicht: (2024)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
von: Pan, Jing, et al.
Veröffentlicht: (2023)
von: Pan, Jing, et al.
Veröffentlicht: (2023)
Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
von: Ling, Shaoshi, et al.
Veröffentlicht: (2025)
von: Ling, Shaoshi, et al.
Veröffentlicht: (2025)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
von: Zhang, Shiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Shiqi, et al.
Veröffentlicht: (2025)
Towards the Next Frontier in Speech Representation Learning Using Disentanglement
von: Krishna, Varun, et al.
Veröffentlicht: (2024)
von: Krishna, Varun, et al.
Veröffentlicht: (2024)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
von: Ognjen, et al.
Veröffentlicht: (2024)
von: Ognjen, et al.
Veröffentlicht: (2024)
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition
von: Su, Hsuan, et al.
Veröffentlicht: (2024)
von: Su, Hsuan, et al.
Veröffentlicht: (2024)
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
von: Gu, Zijin, et al.
Veröffentlicht: (2025)
von: Gu, Zijin, et al.
Veröffentlicht: (2025)
Continual Learning in Machine Speech Chain Using Gradient Episodic Memory
von: Tyndall, Geoffrey, et al.
Veröffentlicht: (2024)
von: Tyndall, Geoffrey, et al.
Veröffentlicht: (2024)
SyllableLM: Learning Coarse Semantic Units for Speech Language Models
von: Baade, Alan, et al.
Veröffentlicht: (2024)
von: Baade, Alan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
von: Zheng, Haolong, et al.
Veröffentlicht: (2025) -
Active Learning with Task Adaptation Pre-training for Speech Emotion Recognition
von: Li, Dongyuan, et al.
Veröffentlicht: (2024) -
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
von: Bataev, Vladimir
Veröffentlicht: (2025) -
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
von: Yaish, Ofir, et al.
Veröffentlicht: (2025) -
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
von: Storey, Edward, et al.
Veröffentlicht: (2025)