TOAST: Transformer Optimization using Adaptive and Simple Transformations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cannistraci, Irene, Antonelli, Simone, Palumbo, Emanuele, Sutter, Thomas M., Rodolà, Emanuele, Rieck, Bastian, Vogt, Julia E.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916021407842304
author Cannistraci, Irene
Antonelli, Simone
Palumbo, Emanuele
Sutter, Thomas M.
Rodolà, Emanuele
Rieck, Bastian
Vogt, Julia E.
author_facet Cannistraci, Irene
Antonelli, Simone
Palumbo, Emanuele
Sutter, Thomas M.
Rodolà, Emanuele
Rieck, Bastian
Vogt, Julia E.
contents Foundation models achieve state-of-the-art performance across different tasks, but their size and computational demands raise concerns about accessibility and sustainability. Existing efficiency methods often require additional retraining or finetuning, limiting their practicality. Recent findings suggest that deep neural networks exhibit internal representation similarities. While such similarities across different models have been exploited for enabling techniques such as model stitching and merging, intra-network redundancy remains underexplored as a source for efficiency gains. In this paper, we introduce Transformer Optimization using Adaptive and Simple Transformations (TOAST), a framework that exploits these redundancies to approximate entire transformer blocks with lightweight closed-form mappings, such as linear transformations or even the identity function, without any additional training. Across state-of-the-art pretrained vision models (e.g., ViT, DINOv2, DeiT) and datasets ranging from MNIST to ImageNet-1k, TOAST reduces parameters and computation while preserving, and in some cases improving, downstream performance. These results show that large portions of transformer depth can be replaced by trivial functions, opening a new perspective on efficient foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04941
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TOAST: Transformer Optimization using Adaptive and Simple Transformations
Cannistraci, Irene
Antonelli, Simone
Palumbo, Emanuele
Sutter, Thomas M.
Rodolà, Emanuele
Rieck, Bastian
Vogt, Julia E.
Machine Learning
Artificial Intelligence
Foundation models achieve state-of-the-art performance across different tasks, but their size and computational demands raise concerns about accessibility and sustainability. Existing efficiency methods often require additional retraining or finetuning, limiting their practicality. Recent findings suggest that deep neural networks exhibit internal representation similarities. While such similarities across different models have been exploited for enabling techniques such as model stitching and merging, intra-network redundancy remains underexplored as a source for efficiency gains. In this paper, we introduce Transformer Optimization using Adaptive and Simple Transformations (TOAST), a framework that exploits these redundancies to approximate entire transformer blocks with lightweight closed-form mappings, such as linear transformations or even the identity function, without any additional training. Across state-of-the-art pretrained vision models (e.g., ViT, DINOv2, DeiT) and datasets ranging from MNIST to ImageNet-1k, TOAST reduces parameters and computation while preserving, and in some cases improving, downstream performance. These results show that large portions of transformer depth can be replaced by trivial functions, opening a new perspective on efficient foundation models.
title TOAST: Transformer Optimization using Adaptive and Simple Transformations
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.04941