Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nishida, Naoto, Hiraki, Hirotaka, Rekimoto, Jun, Ishiguro, Yoshio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908320299745280
author Nishida, Naoto
Hiraki, Hirotaka
Rekimoto, Jun
Ishiguro, Yoshio
author_facet Nishida, Naoto
Hiraki, Hirotaka
Rekimoto, Jun
Ishiguro, Yoshio
contents Rich-text captions are essential to help communication for Deaf and hard-of-hearing (DHH) people, second-language learners, and those with autism spectrum disorder (ASD). They also preserve nuances when converting speech to text, enhancing the realism of presentation scripts and conversation or speech logs. However, current real-time captioning systems lack the capability to alter text attributes (ex. capitalization, sizes, and fonts) at the word level, hindering the accurate conveyance of speaker intent that is expressed in the tones or intonations of the speech. For example, ''YOU should do this'' tends to be considered as indicating ''You'' as the focus of the sentence, whereas ''You should do THIS'' tends to be ''This'' as the focus. This paper proposes a solution that changes the text decorations at the word level in real time. As a prototype, we developed an application that adjusts word size based on the loudness of each spoken word. Feedback from users implies that this system helped to convey the speaker's intent, offering a more engaging and accessible captioning experience.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
Nishida, Naoto
Hiraki, Hirotaka
Rekimoto, Jun
Ishiguro, Yoshio
Human-Computer Interaction
Multimedia
Sound
Audio and Speech Processing
Rich-text captions are essential to help communication for Deaf and hard-of-hearing (DHH) people, second-language learners, and those with autism spectrum disorder (ASD). They also preserve nuances when converting speech to text, enhancing the realism of presentation scripts and conversation or speech logs. However, current real-time captioning systems lack the capability to alter text attributes (ex. capitalization, sizes, and fonts) at the word level, hindering the accurate conveyance of speaker intent that is expressed in the tones or intonations of the speech. For example, ''YOU should do this'' tends to be considered as indicating ''You'' as the focus of the sentence, whereas ''You should do THIS'' tends to be ''This'' as the focus. This paper proposes a solution that changes the text decorations at the word level in real time. As a prototype, we developed an application that adjusts word size based on the loudness of each spoken word. Feedback from users implies that this system helped to convey the speaker's intent, offering a more engaging and accessible captioning experience.
title Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
topic Human-Computer Interaction
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2504.10849