Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gabetni, Firas, Curci, Giuseppe, Pilzer, Andrea, Roy, Subhankar, Ricci, Elisa, Franchi, Gianni
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908981512896512
author Gabetni, Firas
Curci, Giuseppe
Pilzer, Andrea
Roy, Subhankar
Ricci, Elisa
Franchi, Gianni
author_facet Gabetni, Firas
Curci, Giuseppe
Pilzer, Andrea
Roy, Subhankar
Ricci, Elisa
Franchi, Gianni
contents Uncertainty quantification (UQ) is essential for deploying deep neural networks in safety-critical settings. Although methods like Deep Ensembles achieve strong UQ performance, their high computational and memory costs hinder scalability to large models. We introduce Hydra Ensembles, an efficient transformer-based ensemble that prunes attention heads to create diverse members and merges them via a new multi-head attention with grouped fully-connected layers. This yields a compact model with inference speed close to a single network, matching or surpassing Deep Ensembles in UQ performance without retraining from scratch. We also provide an in-depth analysis of pruning, showing that naive approaches can harm calibration, whereas Hydra Ensembles preserves robust uncertainty. Experiments on image and text classification tasks, with various architectures, show consistent gains over Deep Ensembles. Remarkably, in zero-shot classification on ImageNet-1k, our approach surpasses state of the art methods, even without requiring additional training.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18358
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
Gabetni, Firas
Curci, Giuseppe
Pilzer, Andrea
Roy, Subhankar
Ricci, Elisa
Franchi, Gianni
Machine Learning
Computer Vision and Pattern Recognition
Uncertainty quantification (UQ) is essential for deploying deep neural networks in safety-critical settings. Although methods like Deep Ensembles achieve strong UQ performance, their high computational and memory costs hinder scalability to large models. We introduce Hydra Ensembles, an efficient transformer-based ensemble that prunes attention heads to create diverse members and merges them via a new multi-head attention with grouped fully-connected layers. This yields a compact model with inference speed close to a single network, matching or surpassing Deep Ensembles in UQ performance without retraining from scratch. We also provide an in-depth analysis of pruning, showing that naive approaches can harm calibration, whereas Hydra Ensembles preserves robust uncertainty. Experiments on image and text classification tasks, with various architectures, show consistent gains over Deep Ensembles. Remarkably, in zero-shot classification on ImageNet-1k, our approach surpasses state of the art methods, even without requiring additional training.
title Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.18358