Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Kuan-Po, Yang, Chih-Kai, Fu, Yu-Kuan, Dunbar, Ewan, Lee, Hung-yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2024)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2024)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
Parallel Synthesis for Autoregressive Speech Generation
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
ConSep: a Noise- and Reverberation-Robust Speech Separation Framework by Magnitude Conditioning
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
von: Ong, Michael, et al.
Veröffentlicht: (2024)
von: Ong, Michael, et al.
Veröffentlicht: (2024)
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
von: Li, Mohan, et al.
Veröffentlicht: (2024)
von: Li, Mohan, et al.
Veröffentlicht: (2024)
What do neural networks listen to? Exploring the crucial bands in Speech Enhancement using Sinc-convolution
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
Multi-Utterance Speech Separation and Association Trained on Short Segments
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2024)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2024)
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
von: Flynn, Robert, et al.
Veröffentlicht: (2026)
von: Flynn, Robert, et al.
Veröffentlicht: (2026)
From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
Cross-Utterance Conditioned VAE for Speech Generation
von: Li, Yang, et al.
Veröffentlicht: (2023)
von: Li, Yang, et al.
Veröffentlicht: (2023)
Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
von: Huang, Chien-yu, et al.
Veröffentlicht: (2023)
von: Huang, Chien-yu, et al.
Veröffentlicht: (2023)
Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
How to Learn a New Language? An Efficient Solution for Self-Supervised Learning Models Unseen Languages Adaption in Low-Resource Scenario
von: Wang, Shih-Heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-Heng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025) -
Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2024) -
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024) -
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024) -
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)