AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kawamura, Kazuki, Rekimoto, Jun
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913255164739584
author Kawamura, Kazuki
Rekimoto, Jun
author_facet Kawamura, Kazuki
Rekimoto, Jun
contents Since humans can listen to audio and watch videos at faster speeds than actually observed, we often listen to or watch these pieces of content at higher playback speeds to increase the time efficiency of content comprehension. To further utilize this capability, systems that automatically adjust the playback speed according to the user's condition and the type of content to assist in more efficient comprehension of time-series content have been developed. However, there is still room for these systems to further extend human speed-listening ability by generating speech with playback speed optimized for even finer time units and providing it to humans. In this study, we determine whether humans can hear the optimized speech and propose a system that automatically adjusts playback speed at units as small as phonemes while ensuring speech intelligibility. The system uses the speech recognizer score as a proxy for how well a human can hear a certain unit of speech and maximizes the speech playback speed to the extent that a human can hear. This method can be used to produce fast but intelligible speech. In the evaluation experiment, we compared the speech played back at a constant fast speed and the flexibly speed-up speech generated by the proposed method in a blind test and confirmed that the proposed method produced speech that was easier to listen to.
format Preprint
id arxiv_https___arxiv_org_abs_2403_02938
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
Kawamura, Kazuki
Rekimoto, Jun
Computation and Language
Human-Computer Interaction
Machine Learning
Sound
Audio and Speech Processing
Since humans can listen to audio and watch videos at faster speeds than actually observed, we often listen to or watch these pieces of content at higher playback speeds to increase the time efficiency of content comprehension. To further utilize this capability, systems that automatically adjust the playback speed according to the user's condition and the type of content to assist in more efficient comprehension of time-series content have been developed. However, there is still room for these systems to further extend human speed-listening ability by generating speech with playback speed optimized for even finer time units and providing it to humans. In this study, we determine whether humans can hear the optimized speech and propose a system that automatically adjusts playback speed at units as small as phonemes while ensuring speech intelligibility. The system uses the speech recognizer score as a proxy for how well a human can hear a certain unit of speech and maximizes the speech playback speed to the extent that a human can hear. This method can be used to produce fast but intelligible speech. In the evaluation experiment, we compared the speech played back at a constant fast speed and the flexibly speed-up speech generated by the proposed method in a blind test and confirmed that the proposed method produced speech that was easier to listen to.
title AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
topic Computation and Language
Human-Computer Interaction
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2403.02938