Word Level Timestamp Generation for Automatic Speech Recognition and Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Ke, Puvvada, Krishna, Rastorgueva, Elena, Chen, Zhehuai, Huang, He, Ding, Shuoyang, Dhawan, Kunal, Xu, Hainan, Balam, Jagadeesh, Ginsburg, Boris
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918029211729920
author Hu, Ke
Puvvada, Krishna
Rastorgueva, Elena
Chen, Zhehuai
Huang, He
Ding, Shuoyang
Dhawan, Kunal
Xu, Hainan
Balam, Jagadeesh
Ginsburg, Boris
author_facet Hu, Ke
Puvvada, Krishna
Rastorgueva, Elena
Chen, Zhehuai
Huang, He
Ding, Shuoyang
Dhawan, Kunal
Xu, Hainan
Balam, Jagadeesh
Ginsburg, Boris
contents We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for a variety of downstream tasks such as speech content retrieval and timed subtitles. While traditional hybrid systems and end-to-end (E2E) models may employ external modules for timestamp prediction, our approach eliminates the need for separate alignment mechanisms. By leveraging the NeMo Forced Aligner (NFA) as a teacher model, we generate word-level timestamps and train the Canary model to predict timestamps directly. We introduce a new <|timestamp|> token, enabling the Canary model to predict start and end timestamps for each word. Our method demonstrates precision and recall rates between 80% and 90%, with timestamp prediction errors ranging from 20 to 120 ms across four languages, with minimal WER degradation. Additionally, we extend our system to automatic speech translation (AST) tasks, achieving timestamp prediction errors around 200 milliseconds.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15646
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Word Level Timestamp Generation for Automatic Speech Recognition and Translation
Hu, Ke
Puvvada, Krishna
Rastorgueva, Elena
Chen, Zhehuai
Huang, He
Ding, Shuoyang
Dhawan, Kunal
Xu, Hainan
Balam, Jagadeesh
Ginsburg, Boris
Computation and Language
Sound
Audio and Speech Processing
We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for a variety of downstream tasks such as speech content retrieval and timed subtitles. While traditional hybrid systems and end-to-end (E2E) models may employ external modules for timestamp prediction, our approach eliminates the need for separate alignment mechanisms. By leveraging the NeMo Forced Aligner (NFA) as a teacher model, we generate word-level timestamps and train the Canary model to predict timestamps directly. We introduce a new <|timestamp|> token, enabling the Canary model to predict start and end timestamps for each word. Our method demonstrates precision and recall rates between 80% and 90%, with timestamp prediction errors ranging from 20 to 120 ms across four languages, with minimal WER degradation. Additionally, we extend our system to automatic speech translation (AST) tasks, achieving timestamp prediction errors around 200 milliseconds.
title Word Level Timestamp Generation for Automatic Speech Recognition and Translation
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.15646