InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917963270979584 |
|---|---|
| author | Zhang, Tony Brännvall, Rickard |
| author_facet | Zhang, Tony Brännvall, Rickard |
| contents | This work explores optimizing transformer-based language models by integrating model compression techniques with inhibitor attention, a novel alternative attention mechanism. Inhibitor attention employs Manhattan distances and ReLU activations instead of the matrix multiplications and softmax activation of the conventional scaled dot-product attention. This shift offers potential computational and energy savings while maintaining model effectiveness. We propose further adjustments to improve the inhibitor mechanism's training efficiency and evaluate its performance on the DistilBERT architecture. Our knowledge distillation experiments indicate that the modified inhibitor transformer model can achieve competitive performance on standard NLP benchmarks, including General Language Understanding Evaluation (GLUE) and sentiment analysis tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_15983 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer Zhang, Tony Brännvall, Rickard Computation and Language Artificial Intelligence Machine Learning 68T50 (Primary) 68T07, 68Q32 (Secondary) I.2.6; I.2.7; I.5.1 This work explores optimizing transformer-based language models by integrating model compression techniques with inhibitor attention, a novel alternative attention mechanism. Inhibitor attention employs Manhattan distances and ReLU activations instead of the matrix multiplications and softmax activation of the conventional scaled dot-product attention. This shift offers potential computational and energy savings while maintaining model effectiveness. We propose further adjustments to improve the inhibitor mechanism's training efficiency and evaluate its performance on the DistilBERT architecture. Our knowledge distillation experiments indicate that the modified inhibitor transformer model can achieve competitive performance on standard NLP benchmarks, including General Language Understanding Evaluation (GLUE) and sentiment analysis tasks. |
| title | InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer |
| topic | Computation and Language Artificial Intelligence Machine Learning 68T50 (Primary) 68T07, 68Q32 (Secondary) I.2.6; I.2.7; I.5.1 |
| url | https://arxiv.org/abs/2503.15983 |