SENS-ASR: Semantic Embedding injection in Neural-transducer for Streaming Automatic Speech Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dkhissi, Youness, Vielzeuf, Valentin, Allesiardo, Elys, Larcher, Anthony
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917334604578816
author Dkhissi, Youness
Vielzeuf, Valentin
Allesiardo, Elys
Larcher, Anthony
author_facet Dkhissi, Youness
Vielzeuf, Valentin
Allesiardo, Elys
Larcher, Anthony
contents Many Automatic Speech Recognition (ASR) applications require streaming processing of the audio data. In streaming mode, ASR systems need to start transcribing the input stream before it is complete, i.e., the systems have to process a stream of inputs with a limited (or no) future context. Compared to offline mode, this reduction of the future context degrades the performance of Streaming-ASR systems, especially while working with low-latency constraint. In this work, we present SENS-ASR, an approach to enhance the transcription quality of Streaming-ASR by reinforcing the acoustic information with semantic information. This semantic information is extracted from the available past frame-embeddings by a context module. This module is trained using knowledge distillation from a sentence embedding Language Model fine-tuned on the training dataset transcriptions. Experiments on standard datasets show that SENS-ASR significantly improves the Word Error Rate on small-chunk streaming scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10005
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SENS-ASR: Semantic Embedding injection in Neural-transducer for Streaming Automatic Speech Recognition
Dkhissi, Youness
Vielzeuf, Valentin
Allesiardo, Elys
Larcher, Anthony
Computation and Language
Artificial Intelligence
Many Automatic Speech Recognition (ASR) applications require streaming processing of the audio data. In streaming mode, ASR systems need to start transcribing the input stream before it is complete, i.e., the systems have to process a stream of inputs with a limited (or no) future context. Compared to offline mode, this reduction of the future context degrades the performance of Streaming-ASR systems, especially while working with low-latency constraint. In this work, we present SENS-ASR, an approach to enhance the transcription quality of Streaming-ASR by reinforcing the acoustic information with semantic information. This semantic information is extracted from the available past frame-embeddings by a context module. This module is trained using knowledge distillation from a sentence embedding Language Model fine-tuned on the training dataset transcriptions. Experiments on standard datasets show that SENS-ASR significantly improves the Word Error Rate on small-chunk streaming scenarios.
title SENS-ASR: Semantic Embedding injection in Neural-transducer for Streaming Automatic Speech Recognition
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.10005