All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moriya, Takafumi, Mimura, Masato, Tanaka, Tomohiro, Sato, Hiroshi, Masumura, Ryo, Ogawa, Atsunori
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908707445538816
author Moriya, Takafumi
Mimura, Masato
Tanaka, Tomohiro
Sato, Hiroshi
Masumura, Ryo
Ogawa, Atsunori
author_facet Moriya, Takafumi
Mimura, Masato
Tanaka, Tomohiro
Sato, Hiroshi
Masumura, Ryo
Ogawa, Atsunori
contents This paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist temporal classification (CTC), attention-based encoder-decoder (AED), and Transducer, in both offline and streaming modes. While each ASR architecture offers distinct advantages and trade-offs depending on the application, maintaining separate models for each scenario incurs substantial development and deployment costs. To address this issue, we introduce a multi-mode joiner that enables seamless integration of various ASR modes within a single unified model. Experiments show that All-in-One ASR significantly reduces the total model footprint while matching or even surpassing the recognition performance of individually optimized ASR models. Furthermore, joint decoding leverages the complementary strengths of different ASR modes, yielding additional improvements in recognition accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
Moriya, Takafumi
Mimura, Masato
Tanaka, Tomohiro
Sato, Hiroshi
Masumura, Ryo
Ogawa, Atsunori
Audio and Speech Processing
This paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist temporal classification (CTC), attention-based encoder-decoder (AED), and Transducer, in both offline and streaming modes. While each ASR architecture offers distinct advantages and trade-offs depending on the application, maintaining separate models for each scenario incurs substantial development and deployment costs. To address this issue, we introduce a multi-mode joiner that enables seamless integration of various ASR modes within a single unified model. Experiments show that All-in-One ASR significantly reduces the total model footprint while matching or even surpassing the recognition performance of individually optimized ASR models. Furthermore, joint decoding leverages the complementary strengths of different ASR modes, yielding additional improvements in recognition accuracy.
title All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
topic Audio and Speech Processing
url https://arxiv.org/abs/2512.11543