WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Mohan, Do, Cong-Thanh, Keizer, Simon, Farag, Youmna, Stoyanchev, Svetlana, Doddipatla, Rama |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024)
di: Li, Mohan, et al.
Pubblicazione: (2024)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
di: Yao, Jixun, et al.
Pubblicazione: (2025)
di: Yao, Jixun, et al.
Pubblicazione: (2025)
On the use of Performer and Agent Attention for Spoken Language Identification
di: dhiman, Jitendra Kumar, et al.
Pubblicazione: (2025)
di: dhiman, Jitendra Kumar, et al.
Pubblicazione: (2025)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
di: Wang, Helin, et al.
Pubblicazione: (2024)
di: Wang, Helin, et al.
Pubblicazione: (2024)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
di: Jung, Yeonjoon, et al.
Pubblicazione: (2024)
di: Jung, Yeonjoon, et al.
Pubblicazione: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
di: Gong, Cheng, et al.
Pubblicazione: (2023)
di: Gong, Cheng, et al.
Pubblicazione: (2023)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
di: HU, Shujie, et al.
Pubblicazione: (2025)
di: HU, Shujie, et al.
Pubblicazione: (2025)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
di: Yen, Hao, et al.
Pubblicazione: (2024)
di: Yen, Hao, et al.
Pubblicazione: (2024)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
di: Lu, Haitian, et al.
Pubblicazione: (2025)
di: Lu, Haitian, et al.
Pubblicazione: (2025)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
di: Lee, Myungjin, et al.
Pubblicazione: (2026)
di: Lee, Myungjin, et al.
Pubblicazione: (2026)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
di: Bai, Ye, et al.
Pubblicazione: (2024)
di: Bai, Ye, et al.
Pubblicazione: (2024)
GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
di: Wang, Wenbin, et al.
Pubblicazione: (2024)
di: Wang, Wenbin, et al.
Pubblicazione: (2024)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
di: Vendrame, Katia, et al.
Pubblicazione: (2025)
di: Vendrame, Katia, et al.
Pubblicazione: (2025)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
di: Shon, Suwon, et al.
Pubblicazione: (2024)
di: Shon, Suwon, et al.
Pubblicazione: (2024)
Long-Form Speech Generation with Spoken Language Models
di: Park, Se Jin, et al.
Pubblicazione: (2024)
di: Park, Se Jin, et al.
Pubblicazione: (2024)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
Zero- and Few-shot Sound Event Localization and Detection
di: Shimada, Kazuki, et al.
Pubblicazione: (2023)
di: Shimada, Kazuki, et al.
Pubblicazione: (2023)
Zero-shot Cross-lingual Voice Transfer for TTS
di: Biadsy, Fadi, et al.
Pubblicazione: (2024)
di: Biadsy, Fadi, et al.
Pubblicazione: (2024)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
di: Pan, Yu, et al.
Pubblicazione: (2025)
di: Pan, Yu, et al.
Pubblicazione: (2025)
Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
di: Xu, Xuenan, et al.
Pubblicazione: (2024)
di: Xu, Xuenan, et al.
Pubblicazione: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
di: Liu, Huadai, et al.
Pubblicazione: (2024)
di: Liu, Huadai, et al.
Pubblicazione: (2024)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
di: Shao, Mingchen, et al.
Pubblicazione: (2025)
di: Shao, Mingchen, et al.
Pubblicazione: (2025)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages
di: Ranjan, Rishabh, et al.
Pubblicazione: (2025)
di: Ranjan, Rishabh, et al.
Pubblicazione: (2025)
FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
di: Ferreira, Alef Iury Siqueira, et al.
Pubblicazione: (2025)
di: Ferreira, Alef Iury Siqueira, et al.
Pubblicazione: (2025)
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024) -
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024) -
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
di: Arora, Siddhant, et al.
Pubblicazione: (2024) -
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024) -
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
di: Deng, Keqi, et al.
Pubblicazione: (2024)