Pronunciation Assessment with Multi-modal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Kaiqi, Peng, Linkai, Yang, Nan, Zhou, Shuran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
von: Fu, Li, et al.
Veröffentlicht: (2025)
von: Fu, Li, et al.
Veröffentlicht: (2025)
Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
von: Wang, Ke, et al.
Veröffentlicht: (2025)
von: Wang, Ke, et al.
Veröffentlicht: (2025)
MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
von: Chen, Yu-Wen, et al.
Veröffentlicht: (2023)
von: Chen, Yu-Wen, et al.
Veröffentlicht: (2023)
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment
von: Wang, Ke, et al.
Veröffentlicht: (2025)
von: Wang, Ke, et al.
Veröffentlicht: (2025)
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
von: Liu, Changsong, et al.
Veröffentlicht: (2025)
von: Liu, Changsong, et al.
Veröffentlicht: (2025)
Spoken Language Intelligence of Large Language Models for Language Learning
von: Peng, Linkai, et al.
Veröffentlicht: (2023)
von: Peng, Linkai, et al.
Veröffentlicht: (2023)
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
von: Meng, Hao, et al.
Veröffentlicht: (2026)
von: Meng, Hao, et al.
Veröffentlicht: (2026)
Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss
von: Chao, Fu-An, et al.
Veröffentlicht: (2025)
von: Chao, Fu-An, et al.
Veröffentlicht: (2025)
Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing
von: Cheng, Gaofeng, et al.
Veröffentlicht: (2025)
von: Cheng, Gaofeng, et al.
Veröffentlicht: (2025)
Towards Unsupervised Speech Recognition Without Pronunciation Models
von: Ni, Junrui, et al.
Veröffentlicht: (2024)
von: Ni, Junrui, et al.
Veröffentlicht: (2024)
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
von: Do, Heejin, et al.
Veröffentlicht: (2024)
von: Do, Heejin, et al.
Veröffentlicht: (2024)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
von: Nguyen, Tuan-Nam, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan-Nam, et al.
Veröffentlicht: (2025)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
Segmentation-free Goodness of Pronunciation
von: Cao, Xinwei, et al.
Veröffentlicht: (2025)
von: Cao, Xinwei, et al.
Veröffentlicht: (2025)
UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech
von: Kato, Shuhei
Veröffentlicht: (2025)
von: Kato, Shuhei
Veröffentlicht: (2025)
Multi-stage Large Language Model Correction for Speech Recognition
von: Pu, Jie, et al.
Veröffentlicht: (2023)
von: Pu, Jie, et al.
Veröffentlicht: (2023)
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function
von: Li, Shuhe, et al.
Veröffentlicht: (2025)
von: Li, Shuhe, et al.
Veröffentlicht: (2025)
Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Chain of Correction for Full-text Speech Recognition with Large Language Models
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2025)
Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2023)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2023)
HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
von: Yan, Bi-Cheng, et al.
Veröffentlicht: (2025)
von: Yan, Bi-Cheng, et al.
Veröffentlicht: (2025)
Assessment of L2 Oral Proficiency using Speech Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
von: Dutta, Soumya, et al.
Veröffentlicht: (2023)
von: Dutta, Soumya, et al.
Veröffentlicht: (2023)
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
M3TCM: Multi-modal Multi-task Context Model for Utterance Classification in Motivational Interviews
von: Hossain, Sayed Muddashir, et al.
Veröffentlicht: (2024)
von: Hossain, Sayed Muddashir, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for Spontaneous Speech-Based Suicide Risk Detection
von: Gao, Yifan, et al.
Veröffentlicht: (2025)
von: Gao, Yifan, et al.
Veröffentlicht: (2025)
Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
von: Phan, Nhan, et al.
Veröffentlicht: (2025)
von: Phan, Nhan, et al.
Veröffentlicht: (2025)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
Prompting Large Language Models with Audio for General-Purpose Speech Summarization
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
Large Language Models based ASR Error Correction for Child Conversations
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
Enhancing Speech Large Language Models through Reinforced Behavior Alignment
von: Liu, Yansong, et al.
Veröffentlicht: (2025)
von: Liu, Yansong, et al.
Veröffentlicht: (2025)
Full-text Error Correction for Chinese Speech Recognition with Large Language Model
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
Context and System Fusion in Post-ASR Emotion Recognition with Large Language Models
von: Stepachev, Pavel, et al.
Veröffentlicht: (2024)
von: Stepachev, Pavel, et al.
Veröffentlicht: (2024)
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages
von: Meng, Yangyang, et al.
Veröffentlicht: (2025)
von: Meng, Yangyang, et al.
Veröffentlicht: (2025)
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
von: Fu, Li, et al.
Veröffentlicht: (2025) -
Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
von: Wang, Ke, et al.
Veröffentlicht: (2025) -
MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
von: Chen, Yu-Wen, et al.
Veröffentlicht: (2023) -
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment
von: Wang, Ke, et al.
Veröffentlicht: (2025) -
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
von: Liu, Changsong, et al.
Veröffentlicht: (2025)