Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Seungbae, Lee, Daeun, Stark, Brielle, Han, Jinyoung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Text-Aware Adapter for Few-Shot Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2024)
di: Jung, Youngmoon, et al.
Pubblicazione: (2024)
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement
di: Bae, Jae-Sung, et al.
Pubblicazione: (2025)
di: Bae, Jae-Sung, et al.
Pubblicazione: (2025)
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
di: Liu, Yisi, et al.
Pubblicazione: (2025)
di: Liu, Yisi, et al.
Pubblicazione: (2025)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
di: Guo, Chenxu, et al.
Pubblicazione: (2025)
di: Guo, Chenxu, et al.
Pubblicazione: (2025)
Zero-Shot Text-to-Speech as Golden Speech Generator: A Systematic Framework and its Applicability in Automatic Pronunciation Assessment
di: Lo, Tien-Hong, et al.
Pubblicazione: (2024)
di: Lo, Tien-Hong, et al.
Pubblicazione: (2024)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
di: Zhang, Leying, et al.
Pubblicazione: (2026)
di: Zhang, Leying, et al.
Pubblicazione: (2026)
Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models
di: Kabir, Muhammad Ashad, et al.
Pubblicazione: (2026)
di: Kabir, Muhammad Ashad, et al.
Pubblicazione: (2026)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
di: Kim, Jaeyoung, et al.
Pubblicazione: (2024)
di: Kim, Jaeyoung, et al.
Pubblicazione: (2024)
Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
di: Xie, Jiamin, et al.
Pubblicazione: (2025)
di: Xie, Jiamin, et al.
Pubblicazione: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech
di: Bak, Taejun, et al.
Pubblicazione: (2024)
di: Bak, Taejun, et al.
Pubblicazione: (2024)
Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
di: Kumar, Lokesh, et al.
Pubblicazione: (2026)
di: Kumar, Lokesh, et al.
Pubblicazione: (2026)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
di: Jiang, Yicong, et al.
Pubblicazione: (2024)
di: Jiang, Yicong, et al.
Pubblicazione: (2024)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
di: Deng, Wei, et al.
Pubblicazione: (2025)
di: Deng, Wei, et al.
Pubblicazione: (2025)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
di: Li, Haitao, et al.
Pubblicazione: (2026)
di: Li, Haitao, et al.
Pubblicazione: (2026)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice
di: Zhang, Leying, et al.
Pubblicazione: (2026)
di: Zhang, Leying, et al.
Pubblicazione: (2026)
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis
di: Li, Yingahao Aaron, et al.
Pubblicazione: (2024)
di: Li, Yingahao Aaron, et al.
Pubblicazione: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2025)
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2025)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
di: Choi, Yerin, et al.
Pubblicazione: (2024)
di: Choi, Yerin, et al.
Pubblicazione: (2024)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
di: Pan, Yu, et al.
Pubblicazione: (2024)
di: Pan, Yu, et al.
Pubblicazione: (2024)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
BERT-APC: A Reference-free Framework for Automatic Pitch Correction via Musical Context Inference
di: Kim, Sungjae, et al.
Pubblicazione: (2025)
di: Kim, Sungjae, et al.
Pubblicazione: (2025)
Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention
di: Lee, HyeYoung, et al.
Pubblicazione: (2025)
di: Lee, HyeYoung, et al.
Pubblicazione: (2025)
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
di: Lee, Junhyeok, et al.
Pubblicazione: (2025)
di: Lee, Junhyeok, et al.
Pubblicazione: (2025)
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
di: Song, Yakun, et al.
Pubblicazione: (2024)
di: Song, Yakun, et al.
Pubblicazione: (2024)
Open Source State-Of-the-Art Solution for Romanian Speech Recognition
di: Pirlogeanu, Gabriel, et al.
Pubblicazione: (2025)
di: Pirlogeanu, Gabriel, et al.
Pubblicazione: (2025)
Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations
di: Chen, Jinming, et al.
Pubblicazione: (2025)
di: Chen, Jinming, et al.
Pubblicazione: (2025)
Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition
di: Pokel, Niclas, et al.
Pubblicazione: (2025)
di: Pokel, Niclas, et al.
Pubblicazione: (2025)
Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
di: Jin, Zengrui, et al.
Pubblicazione: (2022)
di: Jin, Zengrui, et al.
Pubblicazione: (2022)
Neural Blind Source Separation and Diarization for Distant Speech Recognition
di: Bando, Yoshiaki, et al.
Pubblicazione: (2024)
di: Bando, Yoshiaki, et al.
Pubblicazione: (2024)
Improving Code-Switching Speech Recognition with TTS Data Augmentation
di: Yeo, Yue Heng, et al.
Pubblicazione: (2026)
di: Yeo, Yue Heng, et al.
Pubblicazione: (2026)
Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization
di: Wu, Yihan, et al.
Pubblicazione: (2024)
di: Wu, Yihan, et al.
Pubblicazione: (2024)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
di: Kim, Nam-Gyu, et al.
Pubblicazione: (2025)
di: Kim, Nam-Gyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Text-Aware Adapter for Few-Shot Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2024) -
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement
di: Bae, Jae-Sung, et al.
Pubblicazione: (2025) -
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
di: Liu, Yisi, et al.
Pubblicazione: (2025) -
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
di: Kim, Taesoo, et al.
Pubblicazione: (2025) -
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)