Out-of-distribution generalisation in spoken language understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Porjazovski, Dejan, Moisio, Anssi, Kurimo, Mikko |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
por: Bijoy, Mehedi Hasan, et al.
Publicado: (2025)
por: Bijoy, Mehedi Hasan, et al.
Publicado: (2025)
Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
por: Phan, Nhan, et al.
Publicado: (2025)
por: Phan, Nhan, et al.
Publicado: (2025)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
por: Okocha, Chibuzor, et al.
Publicado: (2025)
por: Okocha, Chibuzor, et al.
Publicado: (2025)
Audio Dialogues: Dialogues dataset for audio and music understanding
por: Goel, Arushi, et al.
Publicado: (2024)
por: Goel, Arushi, et al.
Publicado: (2024)
Optimizing the role of human evaluation in LLM-based spoken document summarization systems
por: Kroll, Margaret, et al.
Publicado: (2024)
por: Kroll, Margaret, et al.
Publicado: (2024)
Towards continually learning new languages
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
Non-native Children's Automatic Speech Assessment Challenge (NOCASA)
por: Getman, Yaroslav, et al.
Publicado: (2025)
por: Getman, Yaroslav, et al.
Publicado: (2025)
You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish
por: Cumbal, Ronald, et al.
Publicado: (2024)
por: Cumbal, Ronald, et al.
Publicado: (2024)
Multilingual acoustic word embeddings for zero-resource languages
por: Jacobs, Christiaan
Publicado: (2024)
por: Jacobs, Christiaan
Publicado: (2024)
Hanprome: Modified Hangeul for Expression of foreign language pronunciation
por: Kim, Wonchan, et al.
Publicado: (2024)
por: Kim, Wonchan, et al.
Publicado: (2024)
StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
por: Zhang, Yu, et al.
Publicado: (2023)
por: Zhang, Yu, et al.
Publicado: (2023)
Word-wise intonation model for cross-language TTS systems
por: A., Tomilov A., et al.
Publicado: (2024)
por: A., Tomilov A., et al.
Publicado: (2024)
Transferable speech-to-text large language model alignment module
por: Wu, Boyong, et al.
Publicado: (2024)
por: Wu, Boyong, et al.
Publicado: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
por: Uro, Rémi, et al.
Publicado: (2024)
por: Uro, Rémi, et al.
Publicado: (2024)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
por: Lyth, Dan, et al.
Publicado: (2024)
por: Lyth, Dan, et al.
Publicado: (2024)
GLAP: General contrastive audio-text pretraining across domains and languages
por: Dinkel, Heinrich, et al.
Publicado: (2025)
por: Dinkel, Heinrich, et al.
Publicado: (2025)
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
por: Yeo, Eunjung, et al.
Publicado: (2025)
por: Yeo, Eunjung, et al.
Publicado: (2025)
Robustness assessment of large audio language models in multiple-choice evaluation
por: López, Fernando, et al.
Publicado: (2025)
por: López, Fernando, et al.
Publicado: (2025)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
por: Dhakal, Manish, et al.
Publicado: (2024)
por: Dhakal, Manish, et al.
Publicado: (2024)
AfriHuBERT: A self-supervised speech representation model for African languages
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
por: Zhu, Jian, et al.
Publicado: (2023)
por: Zhu, Jian, et al.
Publicado: (2023)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
por: Malisz, Zofia, et al.
Publicado: (2024)
por: Malisz, Zofia, et al.
Publicado: (2024)
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
por: Li, Chia-Yu, et al.
Publicado: (2024)
por: Li, Chia-Yu, et al.
Publicado: (2024)
Mmm whatcha say? Uncovering distal and proximal context effects in first and second-language word perception using psychophysical reverse correlation
por: Tuttösí, Paige, et al.
Publicado: (2024)
por: Tuttösí, Paige, et al.
Publicado: (2024)
Encoding of lexical tone in self-supervised models of spoken language
por: Shen, Gaofei, et al.
Publicado: (2024)
por: Shen, Gaofei, et al.
Publicado: (2024)
One Whisper to Grade Them All
por: Phan, Nhan, et al.
Publicado: (2025)
por: Phan, Nhan, et al.
Publicado: (2025)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
por: Shah, Muhammad A., et al.
Publicado: (2024)
por: Shah, Muhammad A., et al.
Publicado: (2024)
InstructAudio: Unified speech and music generation with natural language instruction
por: Qiang, Chunyu, et al.
Publicado: (2025)
por: Qiang, Chunyu, et al.
Publicado: (2025)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
Convexity-based Pruning of Speech Representation Models
por: Dorszewski, Teresa, et al.
Publicado: (2024)
por: Dorszewski, Teresa, et al.
Publicado: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
por: Cui, Mingyu, et al.
Publicado: (2024)
por: Cui, Mingyu, et al.
Publicado: (2024)
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
por: Lu, Ke-Han, et al.
Publicado: (2024)
por: Lu, Ke-Han, et al.
Publicado: (2024)
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
por: Xu, Jing, et al.
Publicado: (2024)
por: Xu, Jing, et al.
Publicado: (2024)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
por: Kang, Wonjune, et al.
Publicado: (2024)
por: Kang, Wonjune, et al.
Publicado: (2024)
Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
por: Park, ChaeHun, et al.
Publicado: (2024)
por: Park, ChaeHun, et al.
Publicado: (2024)
Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
por: Cui, Ziyun, et al.
Publicado: (2024)
por: Cui, Ziyun, et al.
Publicado: (2024)
LAMA-UT: Language Agnostic Multilingual ASR through Orthography Unification and Language-Specific Transliteration
por: Lee, Sangmin, et al.
Publicado: (2024)
por: Lee, Sangmin, et al.
Publicado: (2024)
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
por: Zhang, Yu, et al.
Publicado: (2024)
por: Zhang, Yu, et al.
Publicado: (2024)
Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured Tokenization
por: Ou, Longshen, et al.
Publicado: (2024)
por: Ou, Longshen, et al.
Publicado: (2024)
Ejemplares similares
-
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
por: Bijoy, Mehedi Hasan, et al.
Publicado: (2025) -
Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
por: Phan, Nhan, et al.
Publicado: (2025) -
Can large audio language models understand child stuttering speech? speech summarization, and source separation
por: Okocha, Chibuzor, et al.
Publicado: (2025) -
Audio Dialogues: Dialogues dataset for audio and music understanding
por: Goel, Arushi, et al.
Publicado: (2024) -
Optimizing the role of human evaluation in LLM-based spoken document summarization systems
por: Kroll, Margaret, et al.
Publicado: (2024)