EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909347099967488 |
|---|---|
| author | de Seyssel, Maureen D'Avirro, Antony Williams, Adina Dupoux, Emmanuel |
| author_facet | de Seyssel, Maureen D'Avirro, Antony Williams, Adina Dupoux, Emmanuel |
| contents | We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech translation. In both cases, the benchmark evaluates the ability of the model to encode emphasis in the speech input and accurately reproduce it in the output, potentially across a change of speaker and language. As part of the evaluation pipeline, we introduce EmphaClass, a new model that classifies emphasis at the frame or word level. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_14069 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models de Seyssel, Maureen D'Avirro, Antony Williams, Adina Dupoux, Emmanuel Computation and Language Sound Audio and Speech Processing We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech translation. In both cases, the benchmark evaluates the ability of the model to encode emphasis in the speech input and accurately reproduce it in the output, potentially across a change of speaker and language. As part of the evaluation pipeline, we introduce EmphaClass, a new model that classifies emphasis at the frame or word level. |
| title | EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2312.14069 |