TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Shashi, Madikeri, Srikanth, Zuluaga-Gomez, Juan, Thorbecke, Iuliia, Villatoro-Tello, Esaú, Burdisso, Sergio, Motlicek, Petr, Pandia, Karthik, Ganapathiraju, Aravind
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909341668343808
author Kumar, Shashi
Madikeri, Srikanth
Zuluaga-Gomez, Juan
Thorbecke, Iuliia
Villatoro-Tello, Esaú
Burdisso, Sergio
Motlicek, Petr
Pandia, Karthik
Ganapathiraju, Aravind
author_facet Kumar, Shashi
Madikeri, Srikanth
Zuluaga-Gomez, Juan
Thorbecke, Iuliia
Villatoro-Tello, Esaú
Burdisso, Sergio
Motlicek, Petr
Pandia, Karthik
Ganapathiraju, Aravind
contents In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent processing with different NLP models for tasks like semantic endpointing and named entity recognition (NER). Our paper introduces TokenVerse, a single Transducer-based model designed to handle multiple tasks. This is achieved by integrating task-specific tokens into the reference text during ASR model training, streamlining the inference and eliminating the need for separate NLP models. In addition to ASR, we conduct experiments on 3 different tasks: speaker change detection, endpointing, and NER. Our experiments on a public and a private dataset show that the proposed method improves ASR by up to 7.7% in relative WER while outperforming the cascaded pipeline approach in individual task performance. Our code is publicly available: https://github.com/idiap/tokenverse-unifying-speech-nlp
format Preprint
id arxiv_https___arxiv_org_abs_2407_04444
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
Kumar, Shashi
Madikeri, Srikanth
Zuluaga-Gomez, Juan
Thorbecke, Iuliia
Villatoro-Tello, Esaú
Burdisso, Sergio
Motlicek, Petr
Pandia, Karthik
Ganapathiraju, Aravind
Computation and Language
Sound
Audio and Speech Processing
In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent processing with different NLP models for tasks like semantic endpointing and named entity recognition (NER). Our paper introduces TokenVerse, a single Transducer-based model designed to handle multiple tasks. This is achieved by integrating task-specific tokens into the reference text during ASR model training, streamlining the inference and eliminating the need for separate NLP models. In addition to ASR, we conduct experiments on 3 different tasks: speaker change detection, endpointing, and NER. Our experiments on a public and a private dataset show that the proposed method improves ASR by up to 7.7% in relative WER while outperforming the cascaded pipeline approach in individual task performance. Our code is publicly available: https://github.com/idiap/tokenverse-unifying-speech-nlp
title TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2407.04444