InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tony, Brännvall, Rickard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917963270979584
author Zhang, Tony
Brännvall, Rickard
author_facet Zhang, Tony
Brännvall, Rickard
contents This work explores optimizing transformer-based language models by integrating model compression techniques with inhibitor attention, a novel alternative attention mechanism. Inhibitor attention employs Manhattan distances and ReLU activations instead of the matrix multiplications and softmax activation of the conventional scaled dot-product attention. This shift offers potential computational and energy savings while maintaining model effectiveness. We propose further adjustments to improve the inhibitor mechanism's training efficiency and evaluate its performance on the DistilBERT architecture. Our knowledge distillation experiments indicate that the modified inhibitor transformer model can achieve competitive performance on standard NLP benchmarks, including General Language Understanding Evaluation (GLUE) and sentiment analysis tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15983
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
Zhang, Tony
Brännvall, Rickard
Computation and Language
Artificial Intelligence
Machine Learning
68T50 (Primary) 68T07, 68Q32 (Secondary)
I.2.6; I.2.7; I.5.1
This work explores optimizing transformer-based language models by integrating model compression techniques with inhibitor attention, a novel alternative attention mechanism. Inhibitor attention employs Manhattan distances and ReLU activations instead of the matrix multiplications and softmax activation of the conventional scaled dot-product attention. This shift offers potential computational and energy savings while maintaining model effectiveness. We propose further adjustments to improve the inhibitor mechanism's training efficiency and evaluate its performance on the DistilBERT architecture. Our knowledge distillation experiments indicate that the modified inhibitor transformer model can achieve competitive performance on standard NLP benchmarks, including General Language Understanding Evaluation (GLUE) and sentiment analysis tasks.
title InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
topic Computation and Language
Artificial Intelligence
Machine Learning
68T50 (Primary) 68T07, 68Q32 (Secondary)
I.2.6; I.2.7; I.5.1
url https://arxiv.org/abs/2503.15983