SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saeki, Takaaki, Maiti, Soumi, Takamichi, Shinnosuke, Watanabe, Shinji, Saruwatari, Hiroshi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914930647629824
author Saeki, Takaaki
Maiti, Soumi
Takamichi, Shinnosuke
Watanabe, Shinji
Saruwatari, Hiroshi
author_facet Saeki, Takaaki
Maiti, Soumi
Takamichi, Shinnosuke
Watanabe, Shinji
Saruwatari, Hiroshi
contents While subjective assessments have been the gold standard for evaluating speech generation, there is a growing need for objective metrics that are highly correlated with human subjective judgments due to their cost efficiency. This paper proposes reference-aware automatic evaluation methods for speech generation inspired by evaluation metrics in natural language processing. The proposed SpeechBERTScore computes the BERTScore for self-supervised dense speech features of the generated and reference speech, which can have different sequential lengths. We also propose SpeechBLEU and SpeechTokenDistance, which are computed on speech discrete tokens. The evaluations on synthesized speech show that our method correlates better with human subjective ratings than mel cepstral distortion and a recent mean opinion score prediction model. Also, they are effective in noisy speech evaluation and have cross-lingual applicability.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16812
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
Saeki, Takaaki
Maiti, Soumi
Takamichi, Shinnosuke
Watanabe, Shinji
Saruwatari, Hiroshi
Sound
Audio and Speech Processing
While subjective assessments have been the gold standard for evaluating speech generation, there is a growing need for objective metrics that are highly correlated with human subjective judgments due to their cost efficiency. This paper proposes reference-aware automatic evaluation methods for speech generation inspired by evaluation metrics in natural language processing. The proposed SpeechBERTScore computes the BERTScore for self-supervised dense speech features of the generated and reference speech, which can have different sequential lengths. We also propose SpeechBLEU and SpeechTokenDistance, which are computed on speech discrete tokens. The evaluations on synthesized speech show that our method correlates better with human subjective ratings than mel cepstral distortion and a recent mean opinion score prediction model. Also, they are effective in noisy speech evaluation and have cross-lingual applicability.
title SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.16812