EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: de Seyssel, Maureen, D'Avirro, Antony, Williams, Adina, Dupoux, Emmanuel
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909347099967488
author de Seyssel, Maureen
D'Avirro, Antony
Williams, Adina
Dupoux, Emmanuel
author_facet de Seyssel, Maureen
D'Avirro, Antony
Williams, Adina
Dupoux, Emmanuel
contents We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech translation. In both cases, the benchmark evaluates the ability of the model to encode emphasis in the speech input and accurately reproduce it in the output, potentially across a change of speaker and language. As part of the evaluation pipeline, we introduce EmphaClass, a new model that classifies emphasis at the frame or word level.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14069
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
de Seyssel, Maureen
D'Avirro, Antony
Williams, Adina
Dupoux, Emmanuel
Computation and Language
Sound
Audio and Speech Processing
We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech translation. In both cases, the benchmark evaluates the ability of the model to encode emphasis in the speech input and accurately reproduce it in the output, potentially across a change of speaker and language. As part of the evaluation pipeline, we introduce EmphaClass, a new model that classifies emphasis at the frame or word level.
title EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2312.14069