UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Yudong, Liu, Xiaokang, zhao, Shaofeng, Su, Rongfeng, Yan, Nan, Wang, Lan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data
di: Yang, Yudong, et al.
Pubblicazione: (2024)
di: Yang, Yudong, et al.
Pubblicazione: (2024)
Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection
di: Yu, Hangbin, et al.
Pubblicazione: (2026)
di: Yu, Hangbin, et al.
Pubblicazione: (2026)
From Speech to Profile: A Protocol-Driven LLM Agent for Psychological Profile Generation
di: Yang, Xingjian, et al.
Pubblicazione: (2026)
di: Yang, Xingjian, et al.
Pubblicazione: (2026)
Deep Speech Synthesis from Multimodal Articulatory Representations
di: Wu, Peter, et al.
Pubblicazione: (2024)
di: Wu, Peter, et al.
Pubblicazione: (2024)
Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics
di: Zhang, Ziqian, et al.
Pubblicazione: (2025)
di: Zhang, Ziqian, et al.
Pubblicazione: (2025)
Probing Human Articulatory Constraints in End-to-End TTS with Reverse and Mismatched Speech-Text Directions
di: Khadse, Parth, et al.
Pubblicazione: (2026)
di: Khadse, Parth, et al.
Pubblicazione: (2026)
SpeechQualityLLM: LLM-Based Multimodal Assessment of Speech Quality
di: Monjur, Mahathir, et al.
Pubblicazione: (2025)
di: Monjur, Mahathir, et al.
Pubblicazione: (2025)
Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features
di: Le, Chenqian, et al.
Pubblicazione: (2026)
di: Le, Chenqian, et al.
Pubblicazione: (2026)
Articulatory Feature Prediction from Surface EMG during Speech Production
di: Lee, Jihwan, et al.
Pubblicazione: (2025)
di: Lee, Jihwan, et al.
Pubblicazione: (2025)
MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
di: Shah, Neil, et al.
Pubblicazione: (2024)
di: Shah, Neil, et al.
Pubblicazione: (2024)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
di: Ma, Ziyang, et al.
Pubblicazione: (2026)
di: Ma, Ziyang, et al.
Pubblicazione: (2026)
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
di: Sun, Haoqin, et al.
Pubblicazione: (2026)
di: Sun, Haoqin, et al.
Pubblicazione: (2026)
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
di: Li, Zehua Kcriss, et al.
Pubblicazione: (2024)
di: Li, Zehua Kcriss, et al.
Pubblicazione: (2024)
Training Articulatory Inversion Models for Interspeaker Consistency
di: McGhee, Charles, et al.
Pubblicazione: (2025)
di: McGhee, Charles, et al.
Pubblicazione: (2025)
Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
di: Yang, Yudong, et al.
Pubblicazione: (2025)
di: Yang, Yudong, et al.
Pubblicazione: (2025)
Tracking Articulatory Dynamics in Speech with a Fixed-Weight BiLSTM-CNN Architecture
di: Pillai, Leena G, et al.
Pubblicazione: (2025)
di: Pillai, Leena G, et al.
Pubblicazione: (2025)
Acoustic to Articulatory Inversion of Speech; Data Driven Approaches, Challenges, Applications, and Future Scope
di: Pillai, Leena G, et al.
Pubblicazione: (2025)
di: Pillai, Leena G, et al.
Pubblicazione: (2025)
SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance
di: Lou, Haowei, et al.
Pubblicazione: (2025)
di: Lou, Haowei, et al.
Pubblicazione: (2025)
Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
di: Wang, Siyin, et al.
Pubblicazione: (2024)
di: Wang, Siyin, et al.
Pubblicazione: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
di: Weise, Tobias, et al.
Pubblicazione: (2024)
di: Weise, Tobias, et al.
Pubblicazione: (2024)
An Efficient Transfer Learning Method Based on Adapter with Local Attributes for Speech Emotion Recognition
di: Song, Haoyu, et al.
Pubblicazione: (2025)
di: Song, Haoyu, et al.
Pubblicazione: (2025)
Articulatory strategy as a source of variation in acoustic vowel dynamics
di: Strycharczuk, Patrycja, et al.
Pubblicazione: (2026)
di: Strycharczuk, Patrycja, et al.
Pubblicazione: (2026)
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
di: Peri, Raghuveer, et al.
Pubblicazione: (2024)
di: Peri, Raghuveer, et al.
Pubblicazione: (2024)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
di: Hu, Shujie, et al.
Pubblicazione: (2024)
di: Hu, Shujie, et al.
Pubblicazione: (2024)
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
di: Zhang, Shaolei, et al.
Pubblicazione: (2025)
di: Zhang, Shaolei, et al.
Pubblicazione: (2025)
Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
Scaling Audio-Text Retrieval with Multimodal Large Language Models
di: Xu, Jilan, et al.
Pubblicazione: (2026)
di: Xu, Jilan, et al.
Pubblicazione: (2026)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
di: Cappellazzo, Umberto, et al.
Pubblicazione: (2025)
di: Cappellazzo, Umberto, et al.
Pubblicazione: (2025)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
Closing the Modality Reasoning Gap for Speech Large Language Models
di: Wang, Chaoren, et al.
Pubblicazione: (2026)
di: Wang, Chaoren, et al.
Pubblicazione: (2026)
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
di: Shao, Mingchen, et al.
Pubblicazione: (2025)
di: Shao, Mingchen, et al.
Pubblicazione: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
di: Guo, Shoutao, et al.
Pubblicazione: (2025)
di: Guo, Shoutao, et al.
Pubblicazione: (2025)
Leveraging Large Language Models for Spontaneous Speech-Based Suicide Risk Detection
di: Gao, Yifan, et al.
Pubblicazione: (2025)
di: Gao, Yifan, et al.
Pubblicazione: (2025)
From Hype to Insight: Rethinking Large Language Model Integration in Visual Speech Recognition
di: Jain, Rishabh, et al.
Pubblicazione: (2025)
di: Jain, Rishabh, et al.
Pubblicazione: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
di: Pan, Yu, et al.
Pubblicazione: (2025)
di: Pan, Yu, et al.
Pubblicazione: (2025)
Personalized Neural Speech Codec
di: Jang, Inseon, et al.
Pubblicazione: (2024)
di: Jang, Inseon, et al.
Pubblicazione: (2024)
Documenti analoghi
-
An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data
di: Yang, Yudong, et al.
Pubblicazione: (2024) -
Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection
di: Yu, Hangbin, et al.
Pubblicazione: (2026) -
From Speech to Profile: A Protocol-Driven LLM Agent for Psychological Profile Generation
di: Yang, Xingjian, et al.
Pubblicazione: (2026) -
Deep Speech Synthesis from Multimodal Articulatory Representations
di: Wu, Peter, et al.
Pubblicazione: (2024) -
Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics
di: Zhang, Ziqian, et al.
Pubblicazione: (2025)