Unified Learnable 2D Convolutional Feature Extraction for ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vieting, Peter, Hilmes, Benedikt, Schlüter, Ralf, Ney, Hermann
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909783570776064
author Vieting, Peter
Hilmes, Benedikt
Schlüter, Ralf
Ney, Hermann
author_facet Vieting, Peter
Hilmes, Benedikt
Schlüter, Ralf
Ney, Hermann
contents Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for different tasks. Yet, many of the existing techniques remain heavily influenced by classical methods. While this inductive bias may ease the system design, our work aims to develop a more generic front-end for feature extraction. Furthermore, we seek to unify the front-end architecture contrasting with existing approaches that apply a composition of several layer topologies originating from different sources. The experiments systematically show how to reduce the influence of existing techniques to achieve a generic front-end. The resulting 2D convolutional front-end is parameter-efficient and suitable for a scenario with limited computational resources unlike large models pre-trained on unlabeled audio. The results demonstrate that this generic unified approach is not only feasible but also matches the performance of existing supervised learnable feature extractors.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unified Learnable 2D Convolutional Feature Extraction for ASR
Vieting, Peter
Hilmes, Benedikt
Schlüter, Ralf
Ney, Hermann
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for different tasks. Yet, many of the existing techniques remain heavily influenced by classical methods. While this inductive bias may ease the system design, our work aims to develop a more generic front-end for feature extraction. Furthermore, we seek to unify the front-end architecture contrasting with existing approaches that apply a composition of several layer topologies originating from different sources. The experiments systematically show how to reduce the influence of existing techniques to achieve a generic front-end. The resulting 2D convolutional front-end is parameter-efficient and suitable for a scenario with limited computational resources unlike large models pre-trained on unlabeled audio. The results demonstrate that this generic unified approach is not only feasible but also matches the performance of existing supervised learnable feature extractors.
title Unified Learnable 2D Convolutional Feature Extraction for ASR
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2509.10031