Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lasbordes, Maxence, Falavigna, Daniele, Brutti, Alessio
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911018820567040
author Lasbordes, Maxence
Falavigna, Daniele
Brutti, Alessio
author_facet Lasbordes, Maxence
Falavigna, Daniele
Brutti, Alessio
contents The ability to dynamically adjust the computational load of neural models during inference in a resource aware manner is crucial for on-device processing scenarios, characterised by limited and time-varying computational resources. Early-exit architectures represent an elegant and effective solution, since they can process the input with a subset of their layers, exiting at intermediate branches (the upmost layers are hence removed from the model). From a different perspective, for automatic speech recognition applications there are memory-efficient neural architectures that apply variable frame rate analysis, through downsampling/upsampling operations in the middle layers, reducing the overall number of operations and improving significantly the performance on well established benchmarks. One example is the Zipformer. However, these architectures lack the modularity necessary to inject early-exit branches. With the aim of improving the performance in early-exit models, we propose introducing parallel layers in the architecture that process downsampled versions of their inputs. % in conjunction with standard processing layers. We show that in this way the speech recognition performance on standard benchmarks significantly improve, at the cost of a small increase in the overall number of model parameters but without affecting the inference time.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18035
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
Lasbordes, Maxence
Falavigna, Daniele
Brutti, Alessio
Computation and Language
Sound
Audio and Speech Processing
68T50 (Primary)
I.2.7; I.5.4
The ability to dynamically adjust the computational load of neural models during inference in a resource aware manner is crucial for on-device processing scenarios, characterised by limited and time-varying computational resources. Early-exit architectures represent an elegant and effective solution, since they can process the input with a subset of their layers, exiting at intermediate branches (the upmost layers are hence removed from the model). From a different perspective, for automatic speech recognition applications there are memory-efficient neural architectures that apply variable frame rate analysis, through downsampling/upsampling operations in the middle layers, reducing the overall number of operations and improving significantly the performance on well established benchmarks. One example is the Zipformer. However, these architectures lack the modularity necessary to inject early-exit branches. With the aim of improving the performance in early-exit models, we propose introducing parallel layers in the architecture that process downsampled versions of their inputs. % in conjunction with standard processing layers. We show that in this way the speech recognition performance on standard benchmarks significantly improve, at the cost of a small increase in the overall number of model parameters but without affecting the inference time.
title Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
topic Computation and Language
Sound
Audio and Speech Processing
68T50 (Primary)
I.2.7; I.5.4
url https://arxiv.org/abs/2506.18035