Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ke, Wei, Wenning, Deng, Yan, He, Lei, Zhao, Sheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912594566053888
author Wang, Ke
Wei, Wenning
Deng, Yan
He, Lei
Zhao, Sheng
author_facet Wang, Ke
Wei, Wenning
Deng, Yan
He, Lei
Zhao, Sheng
contents Automatic Pronunciation Assessment (APA) is critical for Computer-Assisted Language Learning (CALL), requiring evaluation across multiple granularities and aspects. Large Multimodal Models (LMMs) present new opportunities for APA, but their effectiveness in fine-grained assessment remains uncertain. This work investigates fine-tuning LMMs for APA using the Speechocean762 dataset and a private corpus. Fine-tuning significantly outperforms zero-shot settings and achieves competitive results on single-granularity tasks compared to public and commercial systems. The model performs well at word and sentence levels, while phoneme-level assessment remains challenging. We also observe that the Pearson Correlation Coefficient (PCC) reaches 0.9, whereas Spearman's rank Correlation Coefficient (SCC) remains around 0.6, suggesting that SCC better reflects ordinal consistency. These findings highlight both the promise and limitations of LMMs for APA and point to future work on fine-grained modeling and rank-aware evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15701
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
Wang, Ke
Wei, Wenning
Deng, Yan
He, Lei
Zhao, Sheng
Computation and Language
Sound
Audio and Speech Processing
Automatic Pronunciation Assessment (APA) is critical for Computer-Assisted Language Learning (CALL), requiring evaluation across multiple granularities and aspects. Large Multimodal Models (LMMs) present new opportunities for APA, but their effectiveness in fine-grained assessment remains uncertain. This work investigates fine-tuning LMMs for APA using the Speechocean762 dataset and a private corpus. Fine-tuning significantly outperforms zero-shot settings and achieves competitive results on single-granularity tasks compared to public and commercial systems. The model performs well at word and sentence levels, while phoneme-level assessment remains challenging. We also observe that the Pearson Correlation Coefficient (PCC) reaches 0.9, whereas Spearman's rank Correlation Coefficient (SCC) remains around 0.6, suggesting that SCC better reflects ordinal consistency. These findings highlight both the promise and limitations of LMMs for APA and point to future work on fine-grained modeling and rank-aware evaluation.
title Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.15701