Multi-Convformer: Extending Conformer with Multiple Convolution Kernels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prabhu, Darshan, Peng, Yifan, Jyothi, Preethi, Watanabe, Shinji
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913443475357696
author Prabhu, Darshan
Peng, Yifan
Jyothi, Preethi
Watanabe, Shinji
author_facet Prabhu, Darshan
Peng, Yifan
Jyothi, Preethi
Watanabe, Shinji
contents Convolutions have become essential in state-of-the-art end-to-end Automatic Speech Recognition~(ASR) systems due to their efficient modelling of local context. Notably, its use in Conformers has led to superior performance compared to vanilla Transformer-based ASR systems. While components other than the convolution module in the Conformer have been reexamined, altering the convolution module itself has been far less explored. Towards this, we introduce Multi-Convformer that uses multiple convolution kernels within the convolution module of the Conformer in conjunction with gating. This helps in improved modeling of local dependencies at varying granularities. Our model rivals existing Conformer variants such as CgMLP and E-Branchformer in performance, while being more parameter efficient. We empirically compare our approach with Conformer and its variants across four different datasets and three different modelling paradigms and show up to 8% relative word error rate~(WER) improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03718
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Convformer: Extending Conformer with Multiple Convolution Kernels
Prabhu, Darshan
Peng, Yifan
Jyothi, Preethi
Watanabe, Shinji
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
Convolutions have become essential in state-of-the-art end-to-end Automatic Speech Recognition~(ASR) systems due to their efficient modelling of local context. Notably, its use in Conformers has led to superior performance compared to vanilla Transformer-based ASR systems. While components other than the convolution module in the Conformer have been reexamined, altering the convolution module itself has been far less explored. Towards this, we introduce Multi-Convformer that uses multiple convolution kernels within the convolution module of the Conformer in conjunction with gating. This helps in improved modeling of local dependencies at varying granularities. Our model rivals existing Conformer variants such as CgMLP and E-Branchformer in performance, while being more parameter efficient. We empirically compare our approach with Conformer and its variants across four different datasets and three different modelling paradigms and show up to 8% relative word error rate~(WER) improvements.
title Multi-Convformer: Extending Conformer with Multiple Convolution Kernels
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2407.03718