Scaling A Simple Approach to Zero-Shot Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Jinming, Pratap, Vineel, Auli, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
by: Yong, Zheng-Xin, et al.
Published: (2025)
by: Yong, Zheng-Xin, et al.
Published: (2025)
Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
by: Yan, Brian, et al.
Published: (2024)
by: Yan, Brian, et al.
Published: (2024)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
by: Zhu, Han, et al.
Published: (2024)
by: Zhu, Han, et al.
Published: (2024)
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
by: Omnilingual ASR team, et al.
Published: (2025)
by: Omnilingual ASR team, et al.
Published: (2025)
DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
by: Liu, Alexander H., et al.
Published: (2023)
by: Liu, Alexander H., et al.
Published: (2023)
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
by: Lian, Jiachen, et al.
Published: (2023)
by: Lian, Jiachen, et al.
Published: (2023)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
by: Yen, Hao, et al.
Published: (2024)
by: Yen, Hao, et al.
Published: (2024)
Revealing Personality Traits: A New Benchmark Dataset for Explainable Personality Recognition on Dialogues
by: Sun, Lei, et al.
Published: (2024)
by: Sun, Lei, et al.
Published: (2024)
Zero-Shot Action Recognition in Surveillance Videos
by: Pereira, Joao, et al.
Published: (2024)
by: Pereira, Joao, et al.
Published: (2024)
Zero-Shot Text-to-Speech for Vietnamese
by: Vu, Thi, et al.
Published: (2025)
by: Vu, Thi, et al.
Published: (2025)
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
by: Shim, Ryan Soh-Eun, et al.
Published: (2026)
by: Shim, Ryan Soh-Eun, et al.
Published: (2026)
OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs
by: Murzaku, John, et al.
Published: (2025)
by: Murzaku, John, et al.
Published: (2025)
Large Language Models for Zero-Shot Multicultural Name Recognition
by: Phonchai, Thanakorn, et al.
Published: (2025)
by: Phonchai, Thanakorn, et al.
Published: (2025)
A Cooperative Multi-Agent Framework for Zero-Shot Named Entity Recognition
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speech
by: Zheng, Zhisheng, et al.
Published: (2025)
by: Zheng, Zhisheng, et al.
Published: (2025)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
by: Higuchi, Yosuke, et al.
Published: (2023)
by: Higuchi, Yosuke, et al.
Published: (2023)
ViSpeechFormer: A Phonemic Approach for Vietnamese Automatic Speech Recognition
by: Nguyen, Khoa Anh, et al.
Published: (2026)
by: Nguyen, Khoa Anh, et al.
Published: (2026)
Zero-resource Speech Translation and Recognition with LLMs
by: Mundnich, Karel, et al.
Published: (2024)
by: Mundnich, Karel, et al.
Published: (2024)
Contrastive Distillation of Emotion Knowledge from LLMs for Zero-Shot Emotion Recognition
by: Niu, Minxue, et al.
Published: (2025)
by: Niu, Minxue, et al.
Published: (2025)
Self-Improving for Zero-Shot Named Entity Recognition with Large Language Models
by: Xie, Tingyu, et al.
Published: (2023)
by: Xie, Tingyu, et al.
Published: (2023)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
by: Doan, Khai Duy, et al.
Published: (2024)
by: Doan, Khai Duy, et al.
Published: (2024)
FlashSpeech: Efficient Zero-Shot Speech Synthesis
by: Ye, Zhen, et al.
Published: (2024)
by: Ye, Zhen, et al.
Published: (2024)
ReZG: Retrieval-Augmented Zero-Shot Counter Narrative Generation for Hate Speech
by: Jiang, Shuyu, et al.
Published: (2023)
by: Jiang, Shuyu, et al.
Published: (2023)
Zero-Shot Performance Prediction for Probabilistic Scaling Laws
by: Schram, Viktoria, et al.
Published: (2025)
by: Schram, Viktoria, et al.
Published: (2025)
Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation
by: Chaudhury, Rohan, et al.
Published: (2024)
by: Chaudhury, Rohan, et al.
Published: (2024)
A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance
by: Melis, Matteo, et al.
Published: (2025)
by: Melis, Matteo, et al.
Published: (2025)
Zero-Shot Conversational Stance Detection: Dataset and Approaches
by: Ding, Yuzhe, et al.
Published: (2025)
by: Ding, Yuzhe, et al.
Published: (2025)
Entity Decomposition with Filtering: A Zero-Shot Clinical Named Entity Recognition Framework
by: Averly, Reza, et al.
Published: (2024)
by: Averly, Reza, et al.
Published: (2024)
Adaptive Test-Time Scaling for Zero-Shot Respiratory Audio Classification
by: Wang, Tsai-Ning, et al.
Published: (2026)
by: Wang, Tsai-Ning, et al.
Published: (2026)
VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text
by: Nguyen, Trieu Hai, et al.
Published: (2025)
by: Nguyen, Trieu Hai, et al.
Published: (2025)
Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation
by: Rahman, Hanif
Published: (2026)
by: Rahman, Hanif
Published: (2026)
SAM-NER: Semantic Archetype Mediation for Zero-Shot Named Entity Recognition
by: Cai, Ruichu, et al.
Published: (2026)
by: Cai, Ruichu, et al.
Published: (2026)
Large-Scale Label Interpretation Learning for Few-Shot Named Entity Recognition
by: Golde, Jonas, et al.
Published: (2024)
by: Golde, Jonas, et al.
Published: (2024)
Pre-Finetuning for Few-Shot Emotional Speech Recognition
by: Chen, Maximillian, et al.
Published: (2023)
by: Chen, Maximillian, et al.
Published: (2023)
Continual Speech Learning with Fused Speech Features
by: Wang, Guitao, et al.
Published: (2025)
by: Wang, Guitao, et al.
Published: (2025)
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
by: Casanova, Edresson, et al.
Published: (2024)
by: Casanova, Edresson, et al.
Published: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
by: Hu, Yuchen, et al.
Published: (2024)
by: Hu, Yuchen, et al.
Published: (2024)
LELA: an LLM-based Entity Linking Approach with Zero-Shot Domain Adaptation
by: Haffoudhi, Samy, et al.
Published: (2026)
by: Haffoudhi, Samy, et al.
Published: (2026)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
by: Peng, Puyuan, et al.
Published: (2024)
by: Peng, Puyuan, et al.
Published: (2024)
llmNER: (Zero|Few)-Shot Named Entity Recognition, Exploiting the Power of Large Language Models
by: Villena, Fabián, et al.
Published: (2024)
by: Villena, Fabián, et al.
Published: (2024)
Similar Items
-
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
by: Yong, Zheng-Xin, et al.
Published: (2025) -
Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
by: Yan, Brian, et al.
Published: (2024) -
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
by: Zhu, Han, et al.
Published: (2024) -
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
by: Omnilingual ASR team, et al.
Published: (2025) -
DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
by: Liu, Alexander H., et al.
Published: (2023)