BaldWhisper: Faster Whisper with Head Shearing and Layer Merging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sy, Yaya, Cerisara, Christophe, Illina, Irina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918477683490816
author Sy, Yaya
Cerisara, Christophe
Illina, Irina
author_facet Sy, Yaya
Cerisara, Christophe
Illina, Irina
contents Pruning large pre-trained transformers in a data-scarce scenario is challenging, as it often requires massive retraining data to recover performance. For instance, Distill-Whisper prunes Whisper by 40 and retrains on 21,000 hours of speech, far beyond what is available for most languages. Can Whisper be made lighter and faster for edge devices in data-scarce settings? Focusing on Bambara with only 32h of speech-to-text data, we propose a new pruning recipe. Instead of vocabulary pruning, which is unsuitable due to frequent code-switching by Bambara speakers, we compress the embeddings with low-rank decomposition and feature distillation. Rather than removing layers, we merge them to limit performance loss. The final model preserves 90 of the original performance while being 48 smaller and 2.15x faster on a MacBook Air M1.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08599
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
Sy, Yaya
Cerisara, Christophe
Illina, Irina
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
Pruning large pre-trained transformers in a data-scarce scenario is challenging, as it often requires massive retraining data to recover performance. For instance, Distill-Whisper prunes Whisper by 40 and retrains on 21,000 hours of speech, far beyond what is available for most languages. Can Whisper be made lighter and faster for edge devices in data-scarce settings? Focusing on Bambara with only 32h of speech-to-text data, we propose a new pruning recipe. Instead of vocabulary pruning, which is unsuitable due to frequent code-switching by Bambara speakers, we compress the embeddings with low-rank decomposition and feature distillation. Rather than removing layers, we merge them to limit performance loss. The final model preserves 90 of the original performance while being 48 smaller and 2.15x faster on a MacBook Air M1.
title BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
url https://arxiv.org/abs/2510.08599