Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Deng, Keqi, Sun, Guangzhi, Woodland, Philip C. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
por: Deng, Keqi, et al.
Publicado: (2025)
por: Deng, Keqi, et al.
Publicado: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
por: Sun, Guangzhi, et al.
Publicado: (2022)
por: Sun, Guangzhi, et al.
Publicado: (2022)
Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
por: Lashkarashvili, Nineli, et al.
Publicado: (2024)
por: Lashkarashvili, Nineli, et al.
Publicado: (2024)
Speaker Adaptation for Quantised End-to-End ASR Models
por: Zhao, Qiuming, et al.
Publicado: (2024)
por: Zhao, Qiuming, et al.
Publicado: (2024)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
por: Zhao, Qiuming, et al.
Publicado: (2024)
por: Zhao, Qiuming, et al.
Publicado: (2024)
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
por: Deng, Keqi, et al.
Publicado: (2023)
por: Deng, Keqi, et al.
Publicado: (2023)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
por: Chen, Junyang, et al.
Publicado: (2026)
por: Chen, Junyang, et al.
Publicado: (2026)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
por: Wang, Zhichao, et al.
Publicado: (2024)
por: Wang, Zhichao, et al.
Publicado: (2024)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
por: Gao, Xiaoxue, et al.
Publicado: (2025)
por: Gao, Xiaoxue, et al.
Publicado: (2025)
CLEP-DG: Contrastive Learning for Speech Emotion Domain Generalization via Soft Prompt Tuning
por: Shi, Jiacheng, et al.
Publicado: (2025)
por: Shi, Jiacheng, et al.
Publicado: (2025)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
por: Higuchi, Yosuke, et al.
Publicado: (2023)
por: Higuchi, Yosuke, et al.
Publicado: (2023)
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
por: Deng, Keqi, et al.
Publicado: (2024)
por: Deng, Keqi, et al.
Publicado: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
por: Jiang, Ziyue, et al.
Publicado: (2023)
por: Jiang, Ziyue, et al.
Publicado: (2023)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
por: Kounadis-Bastian, Dionyssos, et al.
Publicado: (2024)
por: Kounadis-Bastian, Dionyssos, et al.
Publicado: (2024)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
por: Guimarães, Heitor R., et al.
Publicado: (2024)
por: Guimarães, Heitor R., et al.
Publicado: (2024)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
por: Kang, Wonjune, et al.
Publicado: (2022)
por: Kang, Wonjune, et al.
Publicado: (2022)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
por: Qian, Mengjie, et al.
Publicado: (2025)
por: Qian, Mengjie, et al.
Publicado: (2025)
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
por: Yang, Xiaoyu, et al.
Publicado: (2024)
por: Yang, Xiaoyu, et al.
Publicado: (2024)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
por: Zhou, Junzuo, et al.
Publicado: (2024)
por: Zhou, Junzuo, et al.
Publicado: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
por: Ahmad, Hawraz A., et al.
Publicado: (2024)
por: Ahmad, Hawraz A., et al.
Publicado: (2024)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
por: Li, Hanzhao, et al.
Publicado: (2025)
por: Li, Hanzhao, et al.
Publicado: (2025)
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
por: Li, Mohan, et al.
Publicado: (2024)
por: Li, Mohan, et al.
Publicado: (2024)
Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
por: Comanducci, Luca, et al.
Publicado: (2024)
por: Comanducci, Luca, et al.
Publicado: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
por: Yamashita, Natsuo, et al.
Publicado: (2024)
por: Yamashita, Natsuo, et al.
Publicado: (2024)
Zero- and Few-shot Sound Event Localization and Detection
por: Shimada, Kazuki, et al.
Publicado: (2023)
por: Shimada, Kazuki, et al.
Publicado: (2023)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
por: Guo, Yinlin, et al.
Publicado: (2024)
por: Guo, Yinlin, et al.
Publicado: (2024)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
por: Lu, Wenhuan, et al.
Publicado: (2025)
por: Lu, Wenhuan, et al.
Publicado: (2025)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
por: HU, Shujie, et al.
Publicado: (2025)
por: HU, Shujie, et al.
Publicado: (2025)
On Improving Error Resilience of Neural End-to-End Speech Coders
por: Gupta, Kishan, et al.
Publicado: (2024)
por: Gupta, Kishan, et al.
Publicado: (2024)
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
por: Zhuang, Zhuoran, et al.
Publicado: (2026)
por: Zhuang, Zhuoran, et al.
Publicado: (2026)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
por: Xue, Jinlong, et al.
Publicado: (2024)
por: Xue, Jinlong, et al.
Publicado: (2024)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
por: Ghane, Mohsen, et al.
Publicado: (2025)
por: Ghane, Mohsen, et al.
Publicado: (2025)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
por: Lin, Guan-Ting, et al.
Publicado: (2024)
por: Lin, Guan-Ting, et al.
Publicado: (2024)
Unified Pathological Speech Analysis with Prompt Tuning
por: Yang, Fei, et al.
Publicado: (2024)
por: Yang, Fei, et al.
Publicado: (2024)
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
por: Wang, Mengqi, et al.
Publicado: (2025)
por: Wang, Mengqi, et al.
Publicado: (2025)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
por: Yu, Yifeng, et al.
Publicado: (2024)
por: Yu, Yifeng, et al.
Publicado: (2024)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
por: Lei, Shun, et al.
Publicado: (2023)
por: Lei, Shun, et al.
Publicado: (2023)
Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
por: Raimon, Athul, et al.
Publicado: (2024)
por: Raimon, Athul, et al.
Publicado: (2024)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
por: Kushwaha, Saksham Singh, et al.
Publicado: (2024)
por: Kushwaha, Saksham Singh, et al.
Publicado: (2024)
Representation Purification for End-to-End Speech Translation
por: Zhang, Chengwei, et al.
Publicado: (2024)
por: Zhang, Chengwei, et al.
Publicado: (2024)
Ejemplares similares
-
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
por: Deng, Keqi, et al.
Publicado: (2025) -
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
por: Sun, Guangzhi, et al.
Publicado: (2022) -
Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
por: Lashkarashvili, Nineli, et al.
Publicado: (2024) -
Speaker Adaptation for Quantised End-to-End ASR Models
por: Zhao, Qiuming, et al.
Publicado: (2024) -
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
por: Zhao, Qiuming, et al.
Publicado: (2024)