BERT-LID: Leveraging BERT to Improve Spoken Language Identification
Fuente:
arXiv
Salvato in:
| Autori principali: | Nie, Yuting, Zhao, Junhong, Zhang, Wei-Qiang, Bai, Jinfeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
di: Yoon, Ji Won, et al.
Pubblicazione: (2022)
di: Yoon, Ji Won, et al.
Pubblicazione: (2022)
mHuBERT-147: A Compact Multilingual HuBERT Model
di: Boito, Marcely Zanon, et al.
Pubblicazione: (2024)
di: Boito, Marcely Zanon, et al.
Pubblicazione: (2024)
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
di: Bafna, Niyati, et al.
Pubblicazione: (2025)
di: Bafna, Niyati, et al.
Pubblicazione: (2025)
MelHuBERT: A simplified HuBERT on Mel spectrograms
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
di: Li, Jinpeng, et al.
Pubblicazione: (2024)
di: Li, Jinpeng, et al.
Pubblicazione: (2024)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
di: Komatsu, Ryota, et al.
Pubblicazione: (2024)
di: Komatsu, Ryota, et al.
Pubblicazione: (2024)
Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
di: Abdullah, Badr M., et al.
Pubblicazione: (2025)
di: Abdullah, Badr M., et al.
Pubblicazione: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
AudioBERT: Audio Knowledge Augmented Language Model
di: Ok, Hyunjong, et al.
Pubblicazione: (2024)
di: Ok, Hyunjong, et al.
Pubblicazione: (2024)
AfriHuBERT: A self-supervised speech representation model for African languages
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
di: Valente, Martina, et al.
Pubblicazione: (2024)
di: Valente, Martina, et al.
Pubblicazione: (2024)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
Enhancing and Exploring Mild Cognitive Impairment Detection with W2V-BERT-2.0
di: Wang, Yueguan, et al.
Pubblicazione: (2025)
di: Wang, Yueguan, et al.
Pubblicazione: (2025)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
di: Wang, Qingzheng, et al.
Pubblicazione: (2025)
di: Wang, Qingzheng, et al.
Pubblicazione: (2025)
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
di: Poli, Maxime, et al.
Pubblicazione: (2024)
di: Poli, Maxime, et al.
Pubblicazione: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
di: Li, Xingyuan, et al.
Pubblicazione: (2024)
di: Li, Xingyuan, et al.
Pubblicazione: (2024)
On The Landscape of Spoken Language Models: A Comprehensive Survey
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
di: Nakata, Wataru, et al.
Pubblicazione: (2024)
di: Nakata, Wataru, et al.
Pubblicazione: (2024)
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
di: Huo, Robin, et al.
Pubblicazione: (2025)
di: Huo, Robin, et al.
Pubblicazione: (2025)
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
di: Huang, Hukai, et al.
Pubblicazione: (2024)
di: Huang, Hukai, et al.
Pubblicazione: (2024)
Implicit Self-supervised Language Representation for Spoken Language Diarization
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
di: Xie, Jingran, et al.
Pubblicazione: (2025)
di: Xie, Jingran, et al.
Pubblicazione: (2025)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
di: Qian, Mengjie, et al.
Pubblicazione: (2025)
di: Qian, Mengjie, et al.
Pubblicazione: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
di: Lu, Haitian, et al.
Pubblicazione: (2025)
di: Lu, Haitian, et al.
Pubblicazione: (2025)
Spirit LM: Interleaved Spoken and Written Language Model
di: Nguyen, Tu Anh, et al.
Pubblicazione: (2024)
di: Nguyen, Tu Anh, et al.
Pubblicazione: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
Long-Form Speech Generation with Spoken Language Models
di: Park, Se Jin, et al.
Pubblicazione: (2024)
di: Park, Se Jin, et al.
Pubblicazione: (2024)
Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2020)
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2020)
Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
di: Yang, Chih-Kai, et al.
Pubblicazione: (2024)
di: Yang, Chih-Kai, et al.
Pubblicazione: (2024)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
di: Huang, Zhijie, et al.
Pubblicazione: (2026)
di: Huang, Zhijie, et al.
Pubblicazione: (2026)
JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
di: Zhou, Fangru, et al.
Pubblicazione: (2025)
di: Zhou, Fangru, et al.
Pubblicazione: (2025)
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
di: Jung, Yeonjoon, et al.
Pubblicazione: (2024)
di: Jung, Yeonjoon, et al.
Pubblicazione: (2024)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
Evaluating and Improving Continual Learning in Spoken Language Understanding
di: Yang, Muqiao, et al.
Pubblicazione: (2024)
di: Yang, Muqiao, et al.
Pubblicazione: (2024)
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
di: Le, Trang, et al.
Pubblicazione: (2024)
di: Le, Trang, et al.
Pubblicazione: (2024)
Swin-BERT: A Feature Fusion System designed for Speech-based Alzheimer's Dementia Detection
di: Pan, Yilin, et al.
Pubblicazione: (2024)
di: Pan, Yilin, et al.
Pubblicazione: (2024)
"Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
di: Yen, Hao, et al.
Pubblicazione: (2024)
di: Yen, Hao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
di: Yoon, Ji Won, et al.
Pubblicazione: (2022) -
mHuBERT-147: A Compact Multilingual HuBERT Model
di: Boito, Marcely Zanon, et al.
Pubblicazione: (2024) -
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
di: Bafna, Niyati, et al.
Pubblicazione: (2025) -
MelHuBERT: A simplified HuBERT on Mel spectrograms
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022) -
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
di: Li, Jinpeng, et al.
Pubblicazione: (2024)