Target word activity detector: An approach to obtain ASR word boundaries without lexicon

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sivasankaran, Sunit, Sun, Eric, Li, Jinyu, Huang, Yan, Pan, Jing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916405548417024
author Sivasankaran, Sunit
Sun, Eric
Li, Jinyu
Huang, Yan
Pan, Jing
author_facet Sivasankaran, Sunit
Sun, Eric
Li, Jinyu
Huang, Yan
Pan, Jing
contents Obtaining word timestamp information from end-to-end (E2E) ASR models remains challenging due to the lack of explicit time alignment during training. This issue is further complicated in multilingual models. Existing methods, either rely on lexicons or introduce additional tokens, leading to scalability issues and increased computational costs. In this work, we propose a new approach to estimate word boundaries without relying on lexicons. Our method leverages word embeddings from sub-word token units and a pretrained ASR model, requiring only word alignment information during training. Our proposed method can scale-up to any number of languages without incurring any additional cost. We validate our approach using a multilingual ASR model trained on five languages and demonstrate its effectiveness against a strong baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Target word activity detector: An approach to obtain ASR word boundaries without lexicon
Sivasankaran, Sunit
Sun, Eric
Li, Jinyu
Huang, Yan
Pan, Jing
Computation and Language
Sound
Audio and Speech Processing
Obtaining word timestamp information from end-to-end (E2E) ASR models remains challenging due to the lack of explicit time alignment during training. This issue is further complicated in multilingual models. Existing methods, either rely on lexicons or introduce additional tokens, leading to scalability issues and increased computational costs. In this work, we propose a new approach to estimate word boundaries without relying on lexicons. Our method leverages word embeddings from sub-word token units and a pretrained ASR model, requiring only word alignment information during training. Our proposed method can scale-up to any number of languages without incurring any additional cost. We validate our approach using a multilingual ASR model trained on five languages and demonstrate its effectiveness against a strong baseline.
title Target word activity detector: An approach to obtain ASR word boundaries without lexicon
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.13913