Parameter-Efficient Transformer Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ndubuaku, Henry, Talhi, Mouad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912361013575680
author Ndubuaku, Henry
Talhi, Mouad
author_facet Ndubuaku, Henry
Talhi, Mouad
contents Embedding layers in transformer-based NLP models typically account for the largest share of model parameters, scaling with vocabulary size but not yielding performance gains proportional to scale. We propose an alternative approach in which token embedding vectors are first generated deterministically, directly from the token IDs using a Fourier expansion of their normalized values, followed by a lightweight multilayer perceptron (MLP) that captures higher-order interactions. We train standard transformers and our architecture on natural language inference tasks (SNLI and MNLI), and evaluate zero-shot performance on sentence textual similarity (STS-B). Our results demonstrate that the proposed method achieves competitive performance using significantly fewer parameters, trains faster, and operates effectively without the need for dropout. This proof-of-concept study highlights the potential for scalable, memory-efficient language models and motivates further large-scale experimentation based on our findings.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02266
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Parameter-Efficient Transformer Embeddings
Ndubuaku, Henry
Talhi, Mouad
Computation and Language
Artificial Intelligence
Machine Learning
68T07 (Primary) 68T50 (Secondary)
Embedding layers in transformer-based NLP models typically account for the largest share of model parameters, scaling with vocabulary size but not yielding performance gains proportional to scale. We propose an alternative approach in which token embedding vectors are first generated deterministically, directly from the token IDs using a Fourier expansion of their normalized values, followed by a lightweight multilayer perceptron (MLP) that captures higher-order interactions. We train standard transformers and our architecture on natural language inference tasks (SNLI and MNLI), and evaluate zero-shot performance on sentence textual similarity (STS-B). Our results demonstrate that the proposed method achieves competitive performance using significantly fewer parameters, trains faster, and operates effectively without the need for dropout. This proof-of-concept study highlights the potential for scalable, memory-efficient language models and motivates further large-scale experimentation based on our findings.
title Parameter-Efficient Transformer Embeddings
topic Computation and Language
Artificial Intelligence
Machine Learning
68T07 (Primary) 68T50 (Secondary)
url https://arxiv.org/abs/2505.02266